Files
thepraandClaude Opus 5.5 eeb873b816 Load: a flood of fake servers; unique-key lookups no longer scan
tools/pasture/flood/flood.cs answers as twenty fake servers (flood1..20.test)
and sends signed Creates, Likes and Follows at a set rate; load.sh measures
the answers, the queue's wait and processing times, its drain and a persona's
home timeline meanwhile (docs/LOAD.md has the method and the runs).

What the runs found:
- Every unique index was partial on $type: "string", which MongoDB never uses
  for an equality lookup, so every post by ObjectURI, actor by ActorURI,
  deleted object, domain block, remote instance and the rest was a
  collection scan (280 ms a post lookup at 30 000 posts). They are partial on
  $gt: "" now, which an equality on a string implies; MongoDB.Entities
  rebuilds them in place at the next start.
- Two inbox workers capped intake near 110 activities a second:
  Federation:InboxConcurrency and DeliveryConcurrency (default 8) set them.
- The indexes the plan listed as missing: a post's boosts and replies, a
  persona's boosts, who follows an actor, timeline rows by author, a post's
  likes and pins.

At 300 activities a second (200 let through, the rest 429 by the per-origin
limit) the queue wait went from 29 s to 6 ms at p50; with the limits lifted
PrivaPub processes about 900 a second, each in under 10 ms, and the home
timeline stays under 20 ms.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LsXgEaXee4GCU1hwYgPJXw
2026-10-05 19:16:30 +02:00

3.5 KiB

PrivaPub under load

How PrivaPub holds up when the fediverse sends it a lot at once, measured in the pasture with tools/pasture/load.sh against the flood peer, and what each run changed. Numbers are from the workstation (32 cores, PrivaPub and Mongo 8 in podman), so they compare runs with each other; they do not promise a production figure.

The method

  • The crowd: flood (tools/pasture/flood/flood.cs, a .NET file-based app) answers as twenty servers, flood1..20.test, with ten actors each: their actor documents and keys, WebFinger, NodeInfo, and inboxes that accept any Follow with a signed Accept. flood run sends signed activities from random actors to PrivaPub's shared inbox at a set rate: Creates of public notes that mention a persona, Likes of the personas' posts, Follows of them.
  • The audience: load.sh makes the personas load0..N under the pasture root; each follows its own share of the crowd (so a note reaches some homes and not others) and posts twice for the crowd to like.
  • What is measured:
    • the crowd's view: what PrivaPub answered (202, or 429 when an origin is over its limit) and how fast;
    • for every activity PrivaPub took, how long it waited in the job queue and how long processing took (its InteractionEvents);
    • the queue's peak of due jobs, and how long it took to drain once the crowd stopped;
    • a persona's home timeline (/api/v1/timelines/home?limit=40), read four times a second throughout.
  • Each run is printed and kept in tools/pasture/out/load/.

tools/pasture/load.sh --rate=300 --seconds=60 --personas=20 --follows=50 is the reference run below. PrivaPub lets each sending origin burst 300 deliveries and earn back 5 a second (RateLimits:Inbox*), so twenty origins at 300 a second leave about 200 a second through and are answered 429 for the rest, by design.

Runs (2026-10-05)

Run Posts held Queue wait p50 / p95 Processing p50 / p95 Queue peak, drain Home p95
100/s, 2 workers 17 000 12 / 17 ms 11 / 15 ms 4, 0 s 18 ms
300/s, 2 workers 23 000 29 / 47 s 17 / 23 ms 5 104, 49 s 27 ms
300/s, 8 workers 29 000 38 / 60 s 81 / 123 ms 5 786, 63 s 119 ms
300/s, 8 workers, unique indexes fixed 45 000 6 / 9 ms 6 / 8 ms 0, 0 s 14 ms
600/s, origin limits lifted 70 000 7 / 80 ms 6 / 9 ms 0, 0 s 17 ms
1000/s, origin limits lifted 110 000 3.1 / 4.1 s 7 / 11 ms 3 933, 6 s 18 ms

What the runs found:

  • Two inbox workers kept up with about 110 activities a second. Federation:InboxConcurrency and Federation:DeliveryConcurrency (default 8) now set them.
  • More workers made it worse, because every lookup by a unique key scanned its whole collection: those indexes were partial on $type: "string", which MongoDB never uses for an equality lookup. Every post by ObjectURI (each arriving Create checks for a duplicate), every actor by ActorURI, deleted objects, domain blocks, remote instances and the rest. They are partial on $gt: "" now (Indexes.Unique), which an equality on a string implies, and MongoDB.Entities rebuilt them in place at the next start.
  • The indexes the plan listed as missing were added on the way (a post's boosts and replies, a persona's boosts, who follows an actor, a timeline's rows by author, a post's likes and pins).
  • With both fixed, eight workers process about 900 activities a second here, each in under 10 ms, and the home timeline does not notice. Above that the queue grows and drains as soon as the burst ends.