Files
SocialPub/PrivaPub/Domain/Statistics
thepraandClaude Opus 5.5 4f5df743ae M11: the opt-in crawler, PrivaPub-Stargazer, and /stargazing
Off by default (Statistics:Crawler:Enabled), as the owner decided. When it is on:
- CrawlPlan runs hourly, from StatisticsSchedule. It inserts the configured seeds and
  queues up to HostsPerHour servers whose last visit is older than RevisitDays, are not
  paused by the breaker and are not domain-blocked, spread across the hour.
- CrawlInstance visits one server at a time as
  "PrivaPub-Stargazer/<ref> (+<base>/stargazing)":
  - it reads robots.txt (RFC 9309: its own group first, then PrivaPub, then *; longest
    rule wins; a 4xx allows everything; a 5xx or no answer keeps it out);
  - it describes servers that only crawling ever found, through
    InstanceDescriber.Describe with robots.txt as the path filter;
  - it reads /api/v1/instance/peers through the new GetStringArray, which keeps what it
    read from the first 1 MB instead of refusing a large list;
  - it adds the names a server could ever be reached at as "crawled": DNS only,
    punycode, no addresses, ports or hidden services, and the reserved test names only on
    a test network. Never more than MaxNewHostsPerCrawl per visit or MaxHosts in all, and
    never over a touched server.
- It reads nothing but robots.txt, NodeInfo, the instance API and the peers list.
  IFederationHttp.GetText serves robots.txt, and HttpScope.Crawl carries the
  User-Agent.
- /stargazing explains all this and how to keep the crawler out, says whether it is on,
  and credits DB-IP. GET /clientapi/admin/statistics/crawler shows the frontier.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01ELjqpznMFMNrJoJUj6K5p2
2026-10-03 12:19:29 +02:00
..