Files
SocialPub/PrivaPub/Web/Pages/Stargazing.cshtml
T
thepraandClaude Opus 5.5 4f5df743ae M11: the opt-in crawler, PrivaPub-Stargazer, and /stargazing
Off by default (Statistics:Crawler:Enabled), as the owner decided. When it is on:
- CrawlPlan runs hourly, from StatisticsSchedule. It inserts the configured seeds and
  queues up to HostsPerHour servers whose last visit is older than RevisitDays, are not
  paused by the breaker and are not domain-blocked, spread across the hour.
- CrawlInstance visits one server at a time as
  "PrivaPub-Stargazer/<ref> (+<base>/stargazing)":
  - it reads robots.txt (RFC 9309: its own group first, then PrivaPub, then *; longest
    rule wins; a 4xx allows everything; a 5xx or no answer keeps it out);
  - it describes servers that only crawling ever found, through
    InstanceDescriber.Describe with robots.txt as the path filter;
  - it reads /api/v1/instance/peers through the new GetStringArray, which keeps what it
    read from the first 1 MB instead of refusing a large list;
  - it adds the names a server could ever be reached at as "crawled": DNS only,
    punycode, no addresses, ports or hidden services, and the reserved test names only on
    a test network. Never more than MaxNewHostsPerCrawl per visit or MaxHosts in all, and
    never over a touched server.
- It reads nothing but robots.txt, NodeInfo, the instance API and the peers list.
  IFederationHttp.GetText serves robots.txt, and HttpScope.Crawl carries the
  User-Agent.
- /stargazing explains all this and how to keep the crawler out, says whether it is on,
  and credits DB-IP. GET /clientapi/admin/statistics/crawler shows the frontier.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01ELjqpznMFMNrJoJUj6K5p2
2026-10-03 12:19:29 +02:00

49 lines
2.1 KiB
Plaintext

@page "/stargazing"
@model PrivaPub.Web.Pages.StargazingModel
@{
ViewData["Title"] = "Stargazing";
}
<header>
<h1>Stargazing</h1>
<div class="meta">What this server records about other servers, and how to keep its crawler out.</div>
</header>
<article>
<h2>Servers, never people</h2>
<p>This server keeps statistics about the fediverse for teaching and curiosity. They name <strong>servers</strong>,
never accounts: what kind of software a server runs, what it exchanges with us, and how reliably. Distinct accounts are
only counted, through a key that is destroyed at the end of each day.</p>
<p>A server we exchange activities with is described once a week, from its public NodeInfo and, when it has one, its
Mastodon instance API. Its location comes from the address we reached, looked up in an offline database: only the
country is shown publicly for small servers, and only the CDN for servers behind one.</p>
</article>
<article>
<h2>The crawler</h2>
@if (Model.CrawlerEnabled)
{
<p>The crawler is <strong>on</strong> on this server.</p>
}
else
{
<p>The crawler is <strong>off</strong> on this server: it only learns about servers it already exchanges with.</p>
}
<p>When on, it identifies itself as</p>
<p><code>@Model.UserAgent</code></p>
<p>It visits one server a minute, each at most once a week, and reads only:</p>
<ul>
<li><code>/robots.txt</code></li>
<li><code>/.well-known/nodeinfo</code> and the NodeInfo document it points to</li>
<li><code>/api/v2/instance</code> or <code>/api/v1/instance</code></li>
<li><code>/api/v1/instance/peers</code>, to find other servers</li>
</ul>
<p>It never reads accounts, posts, timelines or directories.</p>
<h2>Keeping it out</h2>
<p>Add this to your server's <code>robots.txt</code>:</p>
<pre>User-agent: @StargazerToken
Disallow: /</pre>
<p>If your <code>robots.txt</code> cannot be read because of a server error or a timeout, the crawler stays out too.</p>
</article>
<footer class="meta">IP geolocation by <a href="https://db-ip.com" rel="nofollow noopener noreferrer">DB-IP</a>, CC BY 4.0.</footer>
@functions {
const string StargazerToken = PrivaPub.Federation.Crawler.Stargazer.Token;
}