Off by default (Statistics:Crawler:Enabled), as the owner decided. When it is on:
- CrawlPlan runs hourly, from StatisticsSchedule. It inserts the configured seeds and
queues up to HostsPerHour servers whose last visit is older than RevisitDays, are not
paused by the breaker and are not domain-blocked, spread across the hour.
- CrawlInstance visits one server at a time as
"PrivaPub-Stargazer/<ref> (+<base>/stargazing)":
- it reads robots.txt (RFC 9309: its own group first, then PrivaPub, then *; longest
rule wins; a 4xx allows everything; a 5xx or no answer keeps it out);
- it describes servers that only crawling ever found, through
InstanceDescriber.Describe with robots.txt as the path filter;
- it reads /api/v1/instance/peers through the new GetStringArray, which keeps what it
read from the first 1 MB instead of refusing a large list;
- it adds the names a server could ever be reached at as "crawled": DNS only,
punycode, no addresses, ports or hidden services, and the reserved test names only on
a test network. Never more than MaxNewHostsPerCrawl per visit or MaxHosts in all, and
never over a touched server.
- It reads nothing but robots.txt, NodeInfo, the instance API and the peers list.
IFederationHttp.GetText serves robots.txt, and HttpScope.Crawl carries the
User-Agent.
- /stargazing explains all this and how to keep the crawler out, says whether it is on,
and credits DB-IP. GET /clientapi/admin/statistics/crawler shows the frontier.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01ELjqpznMFMNrJoJUj6K5p2
- HttpScope (AsyncLocal) tags each outbound request with a purpose and a trigger:
- purpose is set by the caller: actor, key, object, webfinger, context, nodeinfo;
- trigger is set by the job kind, by "verify" during inbox verification, or defaults
to "request".
- FederationHttp records every JSON, media and stream fetch: status, time, bytes, hops,
and an outcome of ok, refused or failed, with a reason: disallowed, remembered,
bad-redirect, too-many-redirects, content-type, too-large, bad-json, private-address,
timeout, network, or the status. A fetch a reader caused (trigger "request") is only
counted per server per day.
- Link previews record a 'preview' event: card, no-card or failed.
- The media proxy counts cache hits.
- TrafficMeter counts the client API per endpoint group, method and status class. It
counts our served documents (actor, outbox, collection, object, activity, licence,
webfinger, nodeinfo) by kind, status and whether signed, per day and never per server,
and never names a circle's collections.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01ELjqpznMFMNrJoJUj6K5p2