Summary
cache_release_window_for_streams (graphforge-filesystem/src/cache_io.rs:16) divides a fixed 64 MiB budget by the number of active streams. With 256 partitions plus one that is
64 MiB / 257 = 255 KB per stream
and DurableFileCacheWriter::write calls synchronize_pending on every completed window — a full File::sync_all() followed by fadvise(DONTNEED).
So every spill in every family pays a full fsync per 255 KB written, and the window shrinks as partition count rises. That is the mechanism behind the measured merge_fsync_operations scaling at roughly 62 fsyncs per partition and growing linearly with data at a fixed count.
Scale
At S20 the ingest writes roughly 713 B/edge over 16.8M edges — about 12 GB through 255 KB windows, so on the order of 46,000 fsyncs. At 0.5-2 ms each that is 23-92 seconds of a ~170 second ingest, entirely serial.
Measured supporting evidence: merge_fsync_operations was 35,928 on the redesigned engine at S20 against 10,545 on the baseline, and a 4.2M-edge ingest at 256 partitions recorded 18,193 fsyncs plus 18,844 fadvise calls with 49,490 context switches. Effective cores sit at 0.82-0.85, consistent with a single thread spending roughly 15% of wall blocked on its own durability barriers.
Why this is the throughput lever worth pricing
Ingest has zero intra-process parallelism — verified by grep across append, seal, shape, encode and publish. Adding threads is a large programme. This is a constant-factor serial cost already paid on every run, and it is set by one formula.
It is also the shared root of the RSS/fsync trade recorded in #1439: raising partition count to help memory shrinks the window and multiplies fsyncs; lowering it does the reverse. Decoupling the window from stream count removes that coupling entirely.
The question that has to be answered first
What must be durable before a spill is sealed?
The per-window sync_all is only necessary if a crash between windows must leave the partial spill recoverable. If the correctness bound is "durable before the source is retired" — which is what SpillWriter::seal already establishes, with its own sync_all, root sync and receipt — then the per-window syncs buy nothing and can be replaced by sync_file_range for writeback pacing, or dropped until seal.
This is a durability-policy decision, not a tuning knob. It should be settled explicitly before the formula changes.
Note also that fadvise(DONTNEED) on spills that are read back seconds later converts every logical re-read into a device read. The ladder gate measures process RSS (VmHWM), and no cgroup memory limit is actually enforced — --memlimit appears only in a test fixture, and a baseline rung recorded a 12.41 GB cgroup peak. So the page cache is unconstrained and the eviction may be costing reads for a bound nothing is enforcing.
Acceptance
merge_fsync_operations after the change is proportional to spills plus outputs rather than to bytes written divided by a shrinking window; state the measured counts at two scales.
- Ingest wall time and effective cores at S20, before and after, on a quiet host.
- The crash and recovery suites pass unmodified —
recovery/tests.rs, construction_lifecycle_tests.rs. Prove by fault injection that no spill is retired before it is durable.
- The four fixed-width digests unchanged.
Related
#1439 (the RSS/fsync trade this decouples), #1387, #1417.
Summary
cache_release_window_for_streams(graphforge-filesystem/src/cache_io.rs:16) divides a fixed 64 MiB budget by the number of active streams. With 256 partitions plus one that isand
DurableFileCacheWriter::writecallssynchronize_pendingon every completed window — a fullFile::sync_all()followed byfadvise(DONTNEED).So every spill in every family pays a full fsync per 255 KB written, and the window shrinks as partition count rises. That is the mechanism behind the measured
merge_fsync_operationsscaling at roughly 62 fsyncs per partition and growing linearly with data at a fixed count.Scale
At S20 the ingest writes roughly 713 B/edge over 16.8M edges — about 12 GB through 255 KB windows, so on the order of 46,000 fsyncs. At 0.5-2 ms each that is 23-92 seconds of a ~170 second ingest, entirely serial.
Measured supporting evidence:
merge_fsync_operationswas 35,928 on the redesigned engine at S20 against 10,545 on the baseline, and a 4.2M-edge ingest at 256 partitions recorded 18,193 fsyncs plus 18,844fadvisecalls with 49,490 context switches. Effective cores sit at 0.82-0.85, consistent with a single thread spending roughly 15% of wall blocked on its own durability barriers.Why this is the throughput lever worth pricing
Ingest has zero intra-process parallelism — verified by grep across append, seal, shape, encode and publish. Adding threads is a large programme. This is a constant-factor serial cost already paid on every run, and it is set by one formula.
It is also the shared root of the RSS/fsync trade recorded in #1439: raising partition count to help memory shrinks the window and multiplies fsyncs; lowering it does the reverse. Decoupling the window from stream count removes that coupling entirely.
The question that has to be answered first
What must be durable before a spill is sealed?
The per-window
sync_allis only necessary if a crash between windows must leave the partial spill recoverable. If the correctness bound is "durable before the source is retired" — which is whatSpillWriter::sealalready establishes, with its ownsync_all, root sync and receipt — then the per-window syncs buy nothing and can be replaced bysync_file_rangefor writeback pacing, or dropped until seal.This is a durability-policy decision, not a tuning knob. It should be settled explicitly before the formula changes.
Note also that
fadvise(DONTNEED)on spills that are read back seconds later converts every logical re-read into a device read. The ladder gate measures process RSS (VmHWM), and no cgroup memory limit is actually enforced —--memlimitappears only in a test fixture, and a baseline rung recorded a 12.41 GB cgroup peak. So the page cache is unconstrained and the eviction may be costing reads for a bound nothing is enforcing.Acceptance
merge_fsync_operationsafter the change is proportional to spills plus outputs rather than to bytes written divided by a shrinking window; state the measured counts at two scales.recovery/tests.rs,construction_lifecycle_tests.rs. Prove by fault injection that no spill is retired before it is durable.Related
#1439 (the RSS/fsync trade this decouples), #1387, #1417.