Skip to content

perf(storage): one fsync per 255 KB written, because the cache-release window is divided by stream count #1442

Description

@DecisionNerd

Summary

cache_release_window_for_streams (graphforge-filesystem/src/cache_io.rs:16) divides a fixed 64 MiB budget by the number of active streams. With 256 partitions plus one that is

64 MiB / 257 = 255 KB per stream

and DurableFileCacheWriter::write calls synchronize_pending on every completed window — a full File::sync_all() followed by fadvise(DONTNEED).

So every spill in every family pays a full fsync per 255 KB written, and the window shrinks as partition count rises. That is the mechanism behind the measured merge_fsync_operations scaling at roughly 62 fsyncs per partition and growing linearly with data at a fixed count.

Scale

At S20 the ingest writes roughly 713 B/edge over 16.8M edges — about 12 GB through 255 KB windows, so on the order of 46,000 fsyncs. At 0.5-2 ms each that is 23-92 seconds of a ~170 second ingest, entirely serial.

Measured supporting evidence: merge_fsync_operations was 35,928 on the redesigned engine at S20 against 10,545 on the baseline, and a 4.2M-edge ingest at 256 partitions recorded 18,193 fsyncs plus 18,844 fadvise calls with 49,490 context switches. Effective cores sit at 0.82-0.85, consistent with a single thread spending roughly 15% of wall blocked on its own durability barriers.

Why this is the throughput lever worth pricing

Ingest has zero intra-process parallelism — verified by grep across append, seal, shape, encode and publish. Adding threads is a large programme. This is a constant-factor serial cost already paid on every run, and it is set by one formula.

It is also the shared root of the RSS/fsync trade recorded in #1439: raising partition count to help memory shrinks the window and multiplies fsyncs; lowering it does the reverse. Decoupling the window from stream count removes that coupling entirely.

The question that has to be answered first

What must be durable before a spill is sealed?

The per-window sync_all is only necessary if a crash between windows must leave the partial spill recoverable. If the correctness bound is "durable before the source is retired" — which is what SpillWriter::seal already establishes, with its own sync_all, root sync and receipt — then the per-window syncs buy nothing and can be replaced by sync_file_range for writeback pacing, or dropped until seal.

This is a durability-policy decision, not a tuning knob. It should be settled explicitly before the formula changes.

Note also that fadvise(DONTNEED) on spills that are read back seconds later converts every logical re-read into a device read. The ladder gate measures process RSS (VmHWM), and no cgroup memory limit is actually enforced — --memlimit appears only in a test fixture, and a baseline rung recorded a 12.41 GB cgroup peak. So the page cache is unconstrained and the eviction may be costing reads for a bound nothing is enforcing.

Acceptance

  • merge_fsync_operations after the change is proportional to spills plus outputs rather than to bytes written divided by a shrinking window; state the measured counts at two scales.
  • Ingest wall time and effective cores at S20, before and after, on a quiet host.
  • The crash and recovery suites pass unmodified — recovery/tests.rs, construction_lifecycle_tests.rs. Prove by fault injection that no spill is retired before it is durable.
  • The four fixed-width digests unchanged.

Related

#1439 (the RSS/fsync trade this decouples), #1387, #1417.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions