The public projection of a server's place showed only the country for
servers reporting fewer than 10 users. The owner never decided that
threshold: every located server now shows its city, coordinates (0.1°)
and network, and a CDN-fronted one still shows only its CDN, since the
address reached is the CDN's edge. Statistics:PublicCityMinUsers is
gone; the ROADMAP decision on server locations, CLAUDE.md and the
/stargazing explainer are corrected.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LsXgEaXee4GCU1hwYgPJXw
/api/privapub/v1/instances/:host gains `geo`, the public projection of a
server's place already decided for public server locations
(PublicGeo.Project): city, coordinates and network for servers reporting
at least 10 users and not behind a CDN, the country otherwise, the CDN's
name for a CDN-fronted one, with DB-IP's attribution. `?host[]=`
answers up to 40 servers at once, and this server describes itself:
its host's address located once a day (SelfLocation), or
Statistics:Geo:Self when the owner sets it.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LsXgEaXee4GCU1hwYgPJXw
Owner decisions (2026-10-04, recorded in docs/ROADMAP.md): production runs everything that is built, and nothing waits
on a person running a command.
- Geolocation updates itself. GeoUpdater, a hosted service, checks daily whether each DB-IP Lite database was built this
month. If not, it fetches this month's, or last month's early in the month. It installs a file only once it opens as
the right kind of database, then swaps it in atomically, and the locator reloads at once. Lookups now run under the
lock, so a reload can no longer dispose a reader mid-lookup. The systemd timer, its script and their setup.sh lines
are gone: the root step they needed never happened, and none is needed now. /stargazing names the database in use.
- The admin CLI runs after the app is built, with every service and nothing started.
- `create-root <login> [--admin]` takes the password on stdin; it is how the first login is made while sign-up is
closed.
- `smoke <persona>` keeps the root `deploy-smoke` and an undiscoverable persona, and gives the root a new password
on every run.
- The deploy signs in as @thepra. It runs the CLI, gets a token through the real OAuth flow (tools/smoke/oauth.sh,
moved out of the pasture's privapub_token, which now uses it), checks the signed-in API and that @thepra is
undiscoverable, then revokes the token. PRIVAPUB_SMOKE_TOKEN is gone.
- The deploy also fails when:
- NodeInfo and the instance API disagree about registrations;
- /stargazing does not say the crawler is on;
- the geolocation databases are missing or more than 40 days old.
- The crawler is on in production, seeded with ten large servers of different kinds. FEDERATION.md now describes it
and how to opt out.
- One registrations switch (Registrations:Mode, default Invitations; Open in tests and the pasture). It is read by
open sign-up (403 when closed), NodeInfo `openRegistrations`, and v1 and v2 of the instance API, so they can no longer
disagree. Before, NodeInfo said open and the instance API said closed. Group invitations always work, so
invites_enabled is true.
- A persona edit through /clientapi no longer resets what the Mastodon API set (discoverable, locked, quote policy…):
the theme is merged into the settings instead of replacing them.
650 tests pass. The deploy's smoke step was rehearsed against the pasture's PrivaPub.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01ELjqpznMFMNrJoJUj6K5p2
Off by default (Statistics:Crawler:Enabled), as the owner decided. When it is on:
- CrawlPlan runs hourly, from StatisticsSchedule. It inserts the configured seeds and
queues up to HostsPerHour servers whose last visit is older than RevisitDays, are not
paused by the breaker and are not domain-blocked, spread across the hour.
- CrawlInstance visits one server at a time as
"PrivaPub-Stargazer/<ref> (+<base>/stargazing)":
- it reads robots.txt (RFC 9309: its own group first, then PrivaPub, then *; longest
rule wins; a 4xx allows everything; a 5xx or no answer keeps it out);
- it describes servers that only crawling ever found, through
InstanceDescriber.Describe with robots.txt as the path filter;
- it reads /api/v1/instance/peers through the new GetStringArray, which keeps what it
read from the first 1 MB instead of refusing a large list;
- it adds the names a server could ever be reached at as "crawled": DNS only,
punycode, no addresses, ports or hidden services, and the reserved test names only on
a test network. Never more than MaxNewHostsPerCrawl per visit or MaxHosts in all, and
never over a touched server.
- It reads nothing but robots.txt, NodeInfo, the instance API and the peers list.
IFederationHttp.GetText serves robots.txt, and HttpScope.Crawl carries the
User-Agent.
- /stargazing explains all this and how to keep the crawler out, says whether it is on,
and credits DB-IP. GET /clientapi/admin/statistics/crawler shows the frontier.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01ELjqpznMFMNrJoJUj6K5p2
- Touches: the ledger marks a server as touched when it sends us a verified activity, when
we exchange activities with it, or when we read its actors, keys, objects or WebFinger.
It upserts RemoteInstance.Seen, FirstSeenAt and LastSeenAt at most hourly per server, and
queues one DescribeInstance a week with the same dedupe key ObjectRecords uses. Suspended
servers and pages behind link previews are never described. Migration _010 marks the
servers already known as touched, with their dates.
- InstanceDescriber.Describe(host, crawled, allowed) reads:
- NodeInfo 2.2/2.1/2.0, now with its published user counts, posts, comments,
description, languages and schema version;
- for software with a Mastodon API, /api/v2/instance falling back to v1: title,
languages, registration mode, character limit, API version, source URL.
It never keeps a contact as a field; the raw document is kept for the admin only. It
locates the server from the address our connection reached (DB-IP Lite city and ASN, the
CDN named when fronted) and writes a RemoteInstanceSnapshot per ISO week, unreachable
weeks included. A crawled server is upserted as crawled only on insert, so it never
downgrades a touched one, and robots.txt can deny any path.
- PublicGeo.Project is the only public form of a location: a CDN-fronted server shows its
CDN only, a server reporting at least ten users shows its city, coordinates and network,
any other only its country.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01ELjqpznMFMNrJoJUj6K5p2
A RollupDay job folds each finished day's InteractionEvents into one InstanceDay per server:
- Counters for the admin, keyed channel:activity:object:outcome:reason, plus signature
schemes, audiences, local kinds and features;
- PublicCounters, everything a public page may ever read:
- inbound and outbound activities from an allowlist, on public, unlisted or unaddressed
traffic only, with the outcome collapsed (accepted or dropped, delivered or failed);
- no reasons, no Flag or Block;
- features of public objects;
- health:ok or health:failed from reachability (a 4xx means the server answered);
- latency, wait and byte histograms, and the number of distinct accounts from that day's
hashes.
Re-running a day replaces it and keeps the live Reads counters. The day's salt is then
deleted, so its hashes can never be recomputed, and the next day is queued.
StatisticsSchedule plans today's rollup every hour and catches up any of the last seven
days that have events but no rollup. Waits round up into their bucket.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01ELjqpznMFMNrJoJUj6K5p2
- HttpScope (AsyncLocal) tags each outbound request with a purpose and a trigger:
- purpose is set by the caller: actor, key, object, webfinger, context, nodeinfo;
- trigger is set by the job kind, by "verify" during inbox verification, or defaults
to "request".
- FederationHttp records every JSON, media and stream fetch: status, time, bytes, hops,
and an outcome of ok, refused or failed, with a reason: disallowed, remembered,
bad-redirect, too-many-redirects, content-type, too-large, bad-json, private-address,
timeout, network, or the status. A fetch a reader caused (trigger "request") is only
counted per server per day.
- Link previews record a 'preview' event: card, no-card or failed.
- The media proxy counts cache hits.
- TrafficMeter counts the client API per endpoint group, method and status class. It
counts our served documents (actor, outbox, collection, object, activity, licence,
webfinger, nodeinfo) by kind, status and whether signed, per day and never per server,
and never names a circle's collections.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01ELjqpznMFMNrJoJUj6K5p2
Arrival carries a verdict that handlers set with one line before an existing return
(Arrival.Drop, Reject, Accept, About), so no handler signature changes. InboxProcessor
records it as an 'in' event, with the time taken, the wait since the inbox accepted it, the
attempt, the audience (from the post's visibility, or else the activity's addressing), the
local actor kind, the object's age for updates, deletes and reactions, and the FEP features
of a delivered object. It also records types nothing handles (unknown-type), actors that
cannot be loaded (deferred) and handler failures.
Drop reasons: fetch-failed, unparseable, misattributed, duplicate, deleted, not-addressed,
not-followed, not-public, not-visible, not-deleted, unknown-object, unknown-recipient,
cross-origin, unsupported. Accepted sub-reasons: stored, poll-vote, edit, refresh,
actor-refresh, actor-delete, removed, tombstone-only, auto-accepted, pending,
follow-answer, quote-answer, quote-granted, undone, reaction, reported. Rejected: blocked,
ignored, quote-refused.
A circle's traffic is private and kindless, and a stranger posting into one is just
not-addressed, so the event store cannot reveal a circle.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01ELjqpznMFMNrJoJUj6K5p2
InteractionEvent records one interaction with a remote server: its channel (recv, in, out,
http, preview, crawl), activity and object type, outcome and reason, status, latency, wait,
bytes, attempt, audience, local actor kind, inbox, signature, features and the object's age.
IInteractionLedger.Record never blocks and never throws: events go into a bounded channel
of 10k, a full channel drops and counts, and a hosted service writes batches of up to 1000
every two seconds.
Privacy, as decided by the owner:
- no persona, root, group or activity id, inbox URL, actor URI or sender IP is stored;
- distinct accounts are counted with an HMAC keyed by a per-day salt (InteractionSalt,
upserted so restarts agree, never created for a past day);
- the local actor kind survives only on public and unlisted traffic;
- a host claimed by an unverified sender is kept only if it is already known.
Traffic caused by reading is only counted per day (InstanceDay.Reads, ServerDay). Indexes:
a 90-day TTL on events, unique day rows, and a TTL safety net on salts.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01ELjqpznMFMNrJoJUj6K5p2