Files
SocialPub/PrivaPub/Web/Pages/Stargazing.cshtml
T
thepraandClaude Opus 5.5 436f7da464
Build / Build (push) Successful in 5m11s
Deploy / privapub.thepra.dev (push) Successful in 5m48s
CDNs found by themselves, and servers followed through time
PrivaPub now finds CDNs three ways, best first: the address ranges the
CDNs publish (Cloudflare, Fastly, Amazon CloudFront, Bunny, Gcore,
Imperva), downloaded daily by CdnUpdater and kept in CdnRangeSet; the
CDN's fingerprint in the responses it already gets from a server
(EdgeHintsHandler on the federation client); and the networks that carry
only a CDN. The fixed ASN list is gone; ASNs shared with plain hosting
(AWS, DataPacket) no longer hide a server. A server's Geo records the
CDN, its domain and how it was found, and weekly snapshots now keep the
city and coordinates too.

Servers through time (ServerPlaces): /instances/:host/history lists a
server's weekly snapshots, a CDN-fronted server's geo names the CDN's
domain and where the server was before it (before_cdn), and
/api/privapub/v1/cdns and /cdns/:domain group servers by CDN with week
by week who joined and who left. Owner decisions recorded in ROADMAP.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LsXgEaXee4GCU1hwYgPJXw
2026-10-04 11:33:39 +02:00

60 lines
2.6 KiB
Plaintext

@page "/stargazing"
@model PrivaPub.Web.Pages.StargazingModel
@{
ViewData["Title"] = "Stargazing";
}
<header>
<h1>Stargazing</h1>
<div class="meta">What this server records about other servers, and how to keep its crawler out.</div>
</header>
<article>
<h2>Servers, never people</h2>
<p>This server keeps statistics about the fediverse for teaching and curiosity. They name <strong>servers</strong>,
never accounts: what kind of software a server runs, what it exchanges with us, and how reliably. Distinct accounts are
only counted, through a key that is destroyed at the end of each day.</p>
<p>A server we exchange activities with is described once a week, from its public NodeInfo and, when it has one, its
Mastodon instance API. Its location comes from the address we reached, looked up in an offline database, and is shown
to the city; for a server behind a CDN, only the CDN is shown, since the address is the CDN's. CDNs are recognised from
the address lists they publish and from their marks in the answers we get, and servers are grouped by their CDN week
by week.</p>
@if (Model.GeoSource is { } geo)
{
<p>Locations come from <a href="https://db-ip.com" rel="nofollow noopener">@geo</a>, licensed under
<a href="https://creativecommons.org/licenses/by/4.0/" rel="nofollow noopener">CC BY 4.0</a>.</p>
}
else
{
<p>No location database is loaded yet, so servers are not located.</p>
}
</article>
<article>
<h2>The crawler</h2>
@if (Model.CrawlerEnabled)
{
<p>The crawler is <strong>on</strong> on this server.</p>
}
else
{
<p>The crawler is <strong>off</strong> on this server: it only learns about servers it already exchanges with.</p>
}
<p>When on, it identifies itself as</p>
<p><code>@Model.UserAgent</code></p>
<p>It visits one server a minute, each at most once a week, and reads only:</p>
<ul>
<li><code>/robots.txt</code></li>
<li><code>/.well-known/nodeinfo</code> and the NodeInfo document it points to</li>
<li><code>/api/v2/instance</code> or <code>/api/v1/instance</code></li>
<li><code>/api/v1/instance/peers</code>, to find other servers</li>
</ul>
<p>It never reads accounts, posts, timelines or directories.</p>
<h2>Keeping it out</h2>
<p>Add this to your server's <code>robots.txt</code>:</p>
<pre>User-agent: @StargazerToken
Disallow: /</pre>
<p>If your <code>robots.txt</code> cannot be read because of a server error or a timeout, the crawler stays out too.</p>
</article>
<footer class="meta">IP geolocation by <a href="https://db-ip.com" rel="nofollow noopener noreferrer">DB-IP</a>, CC BY 4.0.</footer>
@functions {
const string StargazerToken = PrivaPub.Federation.Crawler.Stargazer.Token;
}