StacklogBot
This page is what our user agent points at. It is here because anyone whose server we touch has the right to know exactly what we did and how to stop us.
Who is behind this crawler
ARLing s. r. o., Ivanska cesta 32E, 821 04 Bratislava, Slovakia. Company ID 56583486, VAT ID SK2122352100. One person answers the mail: andrej@arling.sk.
The user agent
Every request we make identifies itself and points back at this page.
StacklogBot/1.0 (+https://arling.sk/technologies/crawler/)
What it fetches
- The homepage of the domain, at most 512 KB of HTML, over GET only.
- robots.txt, read before anything else and cached for 24 hours. If it answers with a server error or the connection drops, we fetch nothing at all: an unreadable robots.txt is a refusal, not consent.
- sitemap.xml and the web app manifest, with HEAD where the server allows it.
- One request to a random path that should not exist, to see whether the site answers a real 404.
- Nothing else. No POST, no login, no cookies carried, no admin paths, no /cpanel, no /webmail, no vulnerability probing of any kind.
How often
In the survey, at most once every thirty days per domain. On the live path, at most one fetch of any given domain every sixty seconds across all visitors; everyone else is served the cached result. Across the whole survey we make about two requests a second in total, at night, with at most two retries and a backoff.
Crawl-delay in robots.txt is honoured up to thirty seconds.
What we keep
The findings, their evidence lines, a fingerprint of the response and the log of changes. We do not keep the page content, we do not take screenshots, and we do not store registrant names, e-mail addresses or any other personal data. There is nothing in this pipeline that could become a contact list.
Two ways to stop us, both within 24 hours
Add this to your robots.txt:
User-agent: StacklogBot Disallow: /
Or write one line to andrej@arling.sk with the domain. Either way the domain is removed within 24 hours and the API then answers opted_out instead of a report. No form, no account, no reason required.
Why we obey robots.txt
Not as a courtesy. Measuring public pages in the EU rests on the text and data mining exception in Article 4 of Directive (EU) 2019/790, which applies only where the rightholder has not reserved the use in a machine readable way. robots.txt is that reservation, so obeying it is the basis of our legal position, not a nice gesture.
We also do not extract or reuse a substantial part of anyone else’s database: our numbers come from our own crawl, never from copying someone else’s catalogue.