The ALP-Router crawler
You are probably here because this address showed up in the logs of your site. This page says what the crawler does and how to stop it.
How to recognise it
| User-Agent | ALP-Router/<version> (+https://alp-geo.net/crawler) |
| Comes from | 195.208.21.37, 2a02:408:7722:54:31:177:83:22 |
A request with this User-Agent from any other address is not ours.
What it fetches
/.well-known/alp.jsonof a host, the file in which a site describes itself to ALP routers. Once an hour at most; after a failure the pause doubles, up to a day.- The pages which that file names as its anchors, to check that they say what the file claims. Right after the file changed, then once a day. Half a second at least passes between two pages of one site.
/robots.txt, before the pages.
It follows no links and fetches nothing else. It visits only hosts that were added to the list of the router; it does not discover sites on its own.
What it never does
- It does not fetch a page that
robots.txtcloses to it. Whenrobots.txtanswers with a server error, every page counts as closed. - It does not read more than 2 MB of a page or 1 MB of
alp.json, and does not wait for an answer longer than 10 seconds. - It does not follow a redirect to another host.
- It does not send cookies, fill in forms or run scripts.
How to keep it away
From the pages, with robots.txt:
User-agent: ALP-Router
Disallow: /
From the site altogether: do not publish /.well-known/alp.json. That
file is the invitation; robots.txt is not consulted for it. A host that
does not serve it gets one request for it and is then asked again at growing
intervals, up to once a day.