Summary
PR #1440 cleared the S22 rss_bounded_or_plateaued gate and cost ~23% ingest throughput. Recovering it is worth ~2.40 h at S26 — more than the entire nine-item constant-factor plan combined, and the single largest item on the board.
|
S18 |
S19 |
S20 |
| before #1440 |
99,921 |
102,516 |
101,226 edges/s |
| after #1440 |
80,408 |
78,171 |
77,412 edges/s |
|
-19.5% |
-23.7% |
-23.5% |
S20 wall 319s -> 381s. Projected S26 moves 5.53 h -> 7.93 h against a 3.20 h ceiling.
#1440 must not be reverted. It takes projected S26 ingest RSS from 5.57 GB to 2.17 GB against a 3.44 GB limit (slope 5.08 -> 1.94 B/edge). The memory fix and the throughput loss are believed separable.
Mechanism
#1440 removed the per-partition BufWriter::with_capacity(SPILL_BLOCK_BYTES /* 64 KiB */, hashing) and replaced it with PartitionRun (graph_construction/shape.rs:1056), which accumulates a whole contiguous same-partition run into a Vec<u8> and writes it once. It flushes only on partition change.
The routing key is the record's leading 16 bytes, which is also the staged run's sort key, and the range partition is monotone — so runs really are contiguous. (An earlier "scattered writes" hypothesis was checked and killed.)
Which means: with ~256 partitions and ~33.5M endpoint records at S20, a run averages ~131k records ≈ 4.3 MB. The write buffer went from a fixed 64 KiB (L2-resident) to a multi-megabyte Vec that grows by reallocation, and clear() retains capacity so it stays at peak.
This predicts both symptoms at once: throughput down ~23%, and ingest RSS still growing ~1.94 B/edge after the per-partition buffers were removed.
The fix
Flush PartitionRun when bytes.len() crosses a bound as well as on partition change. Writes stay large enough to amortize the syscall, the buffer stays cache-resident, and the accumulator stops scaling with partition size.
Same shape as the #1442 decision: reprice a bounded constant, do not remove the mechanism. #1440's error was deleting the buffer outright when a bounded one was what it needed.
Acceptance — BOTH, or it is not a fix
- Throughput back to ~100k edges/s at S18/S19/S20.
- Ingest RSS still ~91/97/113 MB with slope near 1.9 B/edge. The memory win is what unblocks S22 and must not regress.
- Report the bound curve (64 KiB / 256 KiB / 1 MiB), not a single value chosen by taste.
- Durable artifacts self-consistent via the determinism suite. TCK green.
Verify first
Histogram PartitionRun.bytes.len() at every flush, per family, at S18. If mean flush size is in the megabytes, confirmed. If flushes are already small this hypothesis is dead and the cost is in DurableFileCacheWriter's per-write window arithmetic instead.
Part of #1387. Follows #1439 / PR #1440.
Summary
PR #1440 cleared the S22
rss_bounded_or_plateauedgate and cost ~23% ingest throughput. Recovering it is worth ~2.40 h at S26 — more than the entire nine-item constant-factor plan combined, and the single largest item on the board.S20 wall 319s -> 381s. Projected S26 moves 5.53 h -> 7.93 h against a 3.20 h ceiling.
#1440 must not be reverted. It takes projected S26 ingest RSS from 5.57 GB to 2.17 GB against a 3.44 GB limit (slope 5.08 -> 1.94 B/edge). The memory fix and the throughput loss are believed separable.
Mechanism
#1440 removed the per-partition
BufWriter::with_capacity(SPILL_BLOCK_BYTES /* 64 KiB */, hashing)and replaced it withPartitionRun(graph_construction/shape.rs:1056), which accumulates a whole contiguous same-partition run into aVec<u8>and writes it once. It flushes only on partition change.The routing key is the record's leading 16 bytes, which is also the staged run's sort key, and the range partition is monotone — so runs really are contiguous. (An earlier "scattered writes" hypothesis was checked and killed.)
Which means: with ~256 partitions and ~33.5M endpoint records at S20, a run averages ~131k records ≈ 4.3 MB. The write buffer went from a fixed 64 KiB (L2-resident) to a multi-megabyte
Vecthat grows by reallocation, andclear()retains capacity so it stays at peak.This predicts both symptoms at once: throughput down ~23%, and ingest RSS still growing ~1.94 B/edge after the per-partition buffers were removed.
The fix
Flush
PartitionRunwhenbytes.len()crosses a bound as well as on partition change. Writes stay large enough to amortize the syscall, the buffer stays cache-resident, and the accumulator stops scaling with partition size.Same shape as the #1442 decision: reprice a bounded constant, do not remove the mechanism. #1440's error was deleting the buffer outright when a bounded one was what it needed.
Acceptance — BOTH, or it is not a fix
Verify first
Histogram
PartitionRun.bytes.len()at every flush, per family, at S18. If mean flush size is in the megabytes, confirmed. If flushes are already small this hypothesis is dead and the cost is inDurableFileCacheWriter's per-write window arithmetic instead.Part of #1387. Follows #1439 / PR #1440.