/api/privapub/v1/instances/:host gains `geo`, the public projection of a
server's place already decided for public server locations
(PublicGeo.Project): city, coordinates and network for servers reporting
at least 10 users and not behind a CDN, the country otherwise, the CDN's
name for a CDN-fronted one, with DB-IP's attribution. `?host[]=`
answers up to 40 servers at once, and this server describes itself:
its host's address located once a day (SelfLocation), or
Statistics:Geo:Self when the owner sets it.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LsXgEaXee4GCU1hwYgPJXw
Off by default (Statistics:Crawler:Enabled), as the owner decided. When it is on:
- CrawlPlan runs hourly, from StatisticsSchedule. It inserts the configured seeds and
queues up to HostsPerHour servers whose last visit is older than RevisitDays, are not
paused by the breaker and are not domain-blocked, spread across the hour.
- CrawlInstance visits one server at a time as
"PrivaPub-Stargazer/<ref> (+<base>/stargazing)":
- it reads robots.txt (RFC 9309: its own group first, then PrivaPub, then *; longest
rule wins; a 4xx allows everything; a 5xx or no answer keeps it out);
- it describes servers that only crawling ever found, through
InstanceDescriber.Describe with robots.txt as the path filter;
- it reads /api/v1/instance/peers through the new GetStringArray, which keeps what it
read from the first 1 MB instead of refusing a large list;
- it adds the names a server could ever be reached at as "crawled": DNS only,
punycode, no addresses, ports or hidden services, and the reserved test names only on
a test network. Never more than MaxNewHostsPerCrawl per visit or MaxHosts in all, and
never over a touched server.
- It reads nothing but robots.txt, NodeInfo, the instance API and the peers list.
IFederationHttp.GetText serves robots.txt, and HttpScope.Crawl carries the
User-Agent.
- /stargazing explains all this and how to keep the crawler out, says whether it is on,
and credits DB-IP. GET /clientapi/admin/statistics/crawler shows the frontier.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01ELjqpznMFMNrJoJUj6K5p2
- Touches: the ledger marks a server as touched when it sends us a verified activity, when
we exchange activities with it, or when we read its actors, keys, objects or WebFinger.
It upserts RemoteInstance.Seen, FirstSeenAt and LastSeenAt at most hourly per server, and
queues one DescribeInstance a week with the same dedupe key ObjectRecords uses. Suspended
servers and pages behind link previews are never described. Migration _010 marks the
servers already known as touched, with their dates.
- InstanceDescriber.Describe(host, crawled, allowed) reads:
- NodeInfo 2.2/2.1/2.0, now with its published user counts, posts, comments,
description, languages and schema version;
- for software with a Mastodon API, /api/v2/instance falling back to v1: title,
languages, registration mode, character limit, API version, source URL.
It never keeps a contact as a field; the raw document is kept for the admin only. It
locates the server from the address our connection reached (DB-IP Lite city and ASN, the
CDN named when fronted) and writes a RemoteInstanceSnapshot per ISO week, unreachable
weeks included. A crawled server is upserted as crawled only on insert, so it never
downgrades a touched one, and robots.txt can deny any path.
- PublicGeo.Project is the only public form of a location: a CDN-fronted server shows its
CDN only, a server reporting at least ten users shows its city, coordinates and network,
any other only its country.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01ELjqpznMFMNrJoJUj6K5p2
Under /clientapi/admin/statistics, admin only (the root JWT's IsAdmin policy, like the domain
blocks; /api tokens are persona tokens):
- GET overview?days: totals per channel, inbound and outbound outcomes, the delivery
success rate, delivery latency p50/p95 from the buckets, active and known servers,
distinct accounts, and the ledger's written/dropped/failed counts;
- GET hosts?days&sort=volume|failures|latency|host&software&page&limit: per server
traffic, refusals, failures, latency, accounts, software and delivery health;
- GET hosts/{host}?days: the server's description, its daily series and its last 100
events;
- GET events?host&channel&outcome&reason&before&limit, and GET server?days for the
ServerDay rows;
- POST rollups/{day} refolds a past day; POST hosts/{host}/describe describes a server
again.
Days not yet rolled up, today included, are folded live from the events, so the numbers are
current. Answers that may carry locations carry the DB-IP attribution.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01ELjqpznMFMNrJoJUj6K5p2