Commit Graph
4 Commits
Author SHA1 Message Date
thepraandClaude Opus 5.5 436f7da464 CDNs found by themselves, and servers followed through time
Build / Build (push) Successful in 5m11s
Deploy / privapub.thepra.dev (push) Successful in 5m48s
PrivaPub now finds CDNs three ways, best first: the address ranges the
CDNs publish (Cloudflare, Fastly, Amazon CloudFront, Bunny, Gcore,
Imperva), downloaded daily by CdnUpdater and kept in CdnRangeSet; the
CDN's fingerprint in the responses it already gets from a server
(EdgeHintsHandler on the federation client); and the networks that carry
only a CDN. The fixed ASN list is gone; ASNs shared with plain hosting
(AWS, DataPacket) no longer hide a server. A server's Geo records the
CDN, its domain and how it was found, and weekly snapshots now keep the
city and coordinates too.

Servers through time (ServerPlaces): /instances/:host/history lists a
server's weekly snapshots, a CDN-fronted server's geo names the CDN's
domain and where the server was before it (before_cdn), and
/api/privapub/v1/cdns and /cdns/:domain group servers by CDN with week
by week who joined and who left. Owner decisions recorded in ROADMAP.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LsXgEaXee4GCU1hwYgPJXw
2026-10-04 11:33:39 +02:00
thepraandClaude Opus 5.5 eb7552e8f3 Server locations: no user threshold
The public projection of a server's place showed only the country for
servers reporting fewer than 10 users. The owner never decided that
threshold: every located server now shows its city, coordinates (0.1°)
and network, and a CDN-fronted one still shows only its CDN, since the
address reached is the CDN's edge. Statistics:PublicCityMinUsers is
gone; the ROADMAP decision on server locations, CLAUDE.md and the
/stargazing explainer are corrected.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LsXgEaXee4GCU1hwYgPJXw
2026-10-04 11:12:02 +02:00
thepraandClaude Opus 5.5 5f56681c01 Everything on, phase 1: geolocation fetches itself, the deploy signs in as @thepra, the crawler is on, sign-up by invitation
Build / Build (push) Successful in 5m1s
Deploy / privapub.thepra.dev (push) Successful in 5m39s
Owner decisions (2026-10-04, recorded in docs/ROADMAP.md): production runs everything that is built, and nothing waits
on a person running a command.

- Geolocation updates itself. GeoUpdater, a hosted service, checks daily whether each DB-IP Lite database was built this
  month. If not, it fetches this month's, or last month's early in the month. It installs a file only once it opens as
  the right kind of database, then swaps it in atomically, and the locator reloads at once. Lookups now run under the
  lock, so a reload can no longer dispose a reader mid-lookup. The systemd timer, its script and their setup.sh lines
  are gone: the root step they needed never happened, and none is needed now. /stargazing names the database in use.
- The admin CLI runs after the app is built, with every service and nothing started.
  - `create-root <login> [--admin]` takes the password on stdin; it is how the first login is made while sign-up is
    closed.
  - `smoke <persona>` keeps the root `deploy-smoke` and an undiscoverable persona, and gives the root a new password
    on every run.
- The deploy signs in as @thepra. It runs the CLI, gets a token through the real OAuth flow (tools/smoke/oauth.sh,
  moved out of the pasture's privapub_token, which now uses it), checks the signed-in API and that @thepra is
  undiscoverable, then revokes the token. PRIVAPUB_SMOKE_TOKEN is gone.
- The deploy also fails when:
  - NodeInfo and the instance API disagree about registrations;
  - /stargazing does not say the crawler is on;
  - the geolocation databases are missing or more than 40 days old.
- The crawler is on in production, seeded with ten large servers of different kinds. FEDERATION.md now describes it
  and how to opt out.
- One registrations switch (Registrations:Mode, default Invitations; Open in tests and the pasture). It is read by
  open sign-up (403 when closed), NodeInfo `openRegistrations`, and v1 and v2 of the instance API, so they can no longer
  disagree. Before, NodeInfo said open and the instance API said closed. Group invitations always work, so
  invites_enabled is true.
- A persona edit through /clientapi no longer resets what the Mastodon API set (discoverable, locked, quote policy…):
  the theme is merged into the settings instead of replacing them.

650 tests pass. The deploy's smoke step was rehearsed against the pasture's PrivaPub.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01ELjqpznMFMNrJoJUj6K5p2
2026-10-04 02:37:38 +02:00
thepraandClaude Opus 5.5 4f5df743ae M11: the opt-in crawler, PrivaPub-Stargazer, and /stargazing
Off by default (Statistics:Crawler:Enabled), as the owner decided. When it is on:
- CrawlPlan runs hourly, from StatisticsSchedule. It inserts the configured seeds and
  queues up to HostsPerHour servers whose last visit is older than RevisitDays, are not
  paused by the breaker and are not domain-blocked, spread across the hour.
- CrawlInstance visits one server at a time as
  "PrivaPub-Stargazer/<ref> (+<base>/stargazing)":
  - it reads robots.txt (RFC 9309: its own group first, then PrivaPub, then *; longest
    rule wins; a 4xx allows everything; a 5xx or no answer keeps it out);
  - it describes servers that only crawling ever found, through
    InstanceDescriber.Describe with robots.txt as the path filter;
  - it reads /api/v1/instance/peers through the new GetStringArray, which keeps what it
    read from the first 1 MB instead of refusing a large list;
  - it adds the names a server could ever be reached at as "crawled": DNS only,
    punycode, no addresses, ports or hidden services, and the reserved test names only on
    a test network. Never more than MaxNewHostsPerCrawl per visit or MaxHosts in all, and
    never over a touched server.
- It reads nothing but robots.txt, NodeInfo, the instance API and the peers list.
  IFederationHttp.GetText serves robots.txt, and HttpScope.Crawl carries the
  User-Agent.
- /stargazing explains all this and how to keep the crawler out, says whether it is on,
  and credits DB-IP. GET /clientapi/admin/statistics/crawler shows the frontier.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01ELjqpznMFMNrJoJUj6K5p2
2026-10-03 12:19:29 +02:00