Skip to content

perf(storage): the PartitionRun accumulator is unbounded and grows with partition size #1445

Description

@DecisionNerd

Summary

PR #1440 cleared the S22 rss_bounded_or_plateaued gate and cost ~23% ingest throughput. Recovering it is worth ~2.40 h at S26 — more than the entire nine-item constant-factor plan combined, and the single largest item on the board.

S18 S19 S20
before #1440 99,921 102,516 101,226 edges/s
after #1440 80,408 78,171 77,412 edges/s
-19.5% -23.7% -23.5%

S20 wall 319s -> 381s. Projected S26 moves 5.53 h -> 7.93 h against a 3.20 h ceiling.

#1440 must not be reverted. It takes projected S26 ingest RSS from 5.57 GB to 2.17 GB against a 3.44 GB limit (slope 5.08 -> 1.94 B/edge). The memory fix and the throughput loss are believed separable.

Mechanism

#1440 removed the per-partition BufWriter::with_capacity(SPILL_BLOCK_BYTES /* 64 KiB */, hashing) and replaced it with PartitionRun (graph_construction/shape.rs:1056), which accumulates a whole contiguous same-partition run into a Vec<u8> and writes it once. It flushes only on partition change.

The routing key is the record's leading 16 bytes, which is also the staged run's sort key, and the range partition is monotone — so runs really are contiguous. (An earlier "scattered writes" hypothesis was checked and killed.)

Which means: with ~256 partitions and ~33.5M endpoint records at S20, a run averages ~131k records ≈ 4.3 MB. The write buffer went from a fixed 64 KiB (L2-resident) to a multi-megabyte Vec that grows by reallocation, and clear() retains capacity so it stays at peak.

This predicts both symptoms at once: throughput down ~23%, and ingest RSS still growing ~1.94 B/edge after the per-partition buffers were removed.

The fix

Flush PartitionRun when bytes.len() crosses a bound as well as on partition change. Writes stay large enough to amortize the syscall, the buffer stays cache-resident, and the accumulator stops scaling with partition size.

Same shape as the #1442 decision: reprice a bounded constant, do not remove the mechanism. #1440's error was deleting the buffer outright when a bounded one was what it needed.

Acceptance — BOTH, or it is not a fix

  • Throughput back to ~100k edges/s at S18/S19/S20.
  • Ingest RSS still ~91/97/113 MB with slope near 1.9 B/edge. The memory win is what unblocks S22 and must not regress.
  • Report the bound curve (64 KiB / 256 KiB / 1 MiB), not a single value chosen by taste.
  • Durable artifacts self-consistent via the determinism suite. TCK green.

Verify first

Histogram PartitionRun.bytes.len() at every flush, per family, at S18. If mean flush size is in the megabytes, confirmed. If flushes are already small this hypothesis is dead and the cost is in DurableFileCacheWriter's per-write window arithmetic instead.

Part of #1387. Follows #1439 / PR #1440.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions