Repository navigation
perf(bench): measure ingest throughput continuously and gate on a floor - #1406
Conversation
GraphForge had no throughput target anywhere and no continuous regression detection on ingest. The existing durable cases cover open, commit, recovery scan, reachability scan, garbage collection and spill compaction; none of them ingests in bulk. The three storage improvements that made the previous evidence stale were each validated by a bespoke one-off measurement, which is why re-running the whole scale ladder was the only way to learn where we stood. Add `ingest_throughput` and `ingest_identifier_density` to the walltime benchmark file, on the isolated CodSpeed Macro Runner alongside the other durable I/O cases, because ingest does real durability work that CPU simulation cannot see. Both stage and publish one generation through the same `GraphConstructionSession` path the scale ladder's ingest phase uses, so the numbers are comparable to the ladder's: 8,388,608 edges publishes at 77,672 edges/sec here against the ladder's 72,989 at 67.1M. Two sizes, not one. The defect is that throughput degrades as the graph grows, and a benchmark at a single fixed size would report a flat healthy number while that continued underneath it. The sweep runs 524,288 and 8,388,608 edges, a sixteen-fold span chosen because below about half a million edges the fixed cost of publishing a generation dominates and the ratio stops describing scaling, while the ladder's 67.1M rung takes fifteen minutes and cannot run nightly. All four ingest benchmarks together cost two minutes. Report bytes read per edge as a first-class metric: 2,139 at the small size and 2,412 at the large one, against 265 bytes retained per edge, the same constant overhead the epic measured as 4,579 at 67.1M. Report CPU microseconds per edge and effective core utilisation too, from `getrusage`: 10.30 and 12.66 µs/edge, 0.44 and 0.59 of 16 cores. The serialized fraction of the path is not measurable from divan or the session's evidence and is not faked. Gate on the floor, not only on regression. CodSpeed compares each run against the previous one, so a slow drift that never regresses in a single step passes indefinitely. `GF_INGEST_FLOOR_GATE` turns the bench binary into a fail-closed gate that runs in the nightly before the measured pass. It fails when throughput at any size drops below a floor, when bytes read or CPU per edge at any size exceed a ceiling, or when bytes read per edge grows more than 1.20x across the sweep. The ratio is evaluated on bytes read per edge because that quantity reproduced to the byte on four runs while wall-clock throughput moved by 48% on the same code; the limits are set at today's measured values and each carries a comment saying it ratchets as the redesign lands. Register both names in the fail-closed inventory, which now also refuses a workflow that does not run the gate, a gate missing any of its four limits, and a sweep collapsed to a single size. Part of #1387 Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
|
Important Review skippedAuto reviews are disabled on this repository. Please check the settings in the CodeRabbit UI or the ⚙️ Run configurationConfiguration used: Path: .coderabbit.yaml Review profile: CHILL Plan: Advanced Run ID: You can disable this status message by setting the Use the checkbox below for a quick retry:
Warning Billing warning: we have not been able to collect payment for this subscription for more than 72 hours. Please update the payment method or pay any pending invoices in Billing to avoid service interruption. Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out. Comment |
…ventory `scripts/ci/test-ci-storage-policy.py` keeps a hand-maintained inventory of every artifact upload name, path and retention across all workflows, so that a new one requires a deliberate entry rather than appearing silently. The ingest floor gate's measurement upload had none, and the Repository Policy job caught it at the moment it was introduced, which is what the inventory is for. Register the name and its path, and conform the upload to the policy it was breaking in two other ways it would have failed on next: `if-no-files-found` must be `error`, not `warn`, and retention is one day for a transfer artifact. Thirty days was wrong — the gate's measurements are neither a publication nor a certification artifact, and their durable record is the CodSpeed walltime series plus the gate step's own log. The gate registry needs nothing: `config/gate-registry.json` is a one-to-one inventory of workflow files and `codspeed.yml` is already registered as the `codspeed` gate. The floor gate is a step inside it, not a new workflow, and `make gate-registry-check` passes unchanged. Part of #1387 Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Workstream 6 of #1387. Ingest had no throughput target anywhere in the project and no continuous regression detection. Every gain the redesign wins without this is unprotected.
What lands
Two divan benchmarks in the walltime file, on the isolated CodSpeed Macro Runner alongside the other durable I/O cases, because ingest does real durability work that CPU simulation cannot see. Both stage and publish one generation through the same
GraphConstructionSessionpath the scale ladder's ingest phase uses, so the numbers are directly comparable to the ladder's.ingest_throughput— 524,288 and 8,388,608 edges at a fixed edge factor of 16.ingest_identifier_density— 131,072 edges at fan-out 2 and 256, the structural axis.Plus a fail-closed floor gate in the same binary, run by the nightly, and both names pinned in the benchmark inventory.
Sizes, and why
A sixteen-fold span, the same shape as the ladder's S18-to-S22 fall.
The lower rung is not smaller because below about half a million edges the fixed cost of opening and publishing a generation dominates and the ratio stops describing scaling: across 131,072 to 2,097,152 edges the read-amplification ratio is 1.046, while across 524,288 to 8,388,608 it is 1.128. The upper rung is capped by the nightly budget, not by what would be most informative — the ladder's 67.1M rung takes about fifteen minutes. All four ingest benchmarks together cost 1m58s, and the gate's own pass costs about three minutes.
Cost of the upper rung: 3.6 GiB of transient construction disk, reported by the gate so the runner requirement stays visible. The job timeout moves 60 → 90 minutes.
Measured
Gate output, this branch, on a 16-core x86 development host:
On an unloaded host the divan pass published 8,388,608 edges in 1.799 min, 77,672 edges/sec, against the ladder's 72,989 at 67.1M and 79,431 at 8.4M. CPU per edge lands at 10.30 to 12.66 µs against the epic's 11.93. The benchmark is measuring the path the epic measured.
The structural axis earns its place: at a constant 131,072 edges, fan-out 256 publishes in 1.2 s and fan-out 2 in 1.885 s, a 1.57x spread from identifier-space density alone that a pure size sweep would have blurred.
The ratio is gated on bytes read per edge
The requirement is to track the ratio between sizes. It is evaluated on bytes read per edge rather than on edges per second, because across four runs of the same commit:
Bytes read per edge reproduced to the byte while wall-clock throughput moved by 48% on identical code. It expresses the same defect — the ladder's throughput fall is read amplification growing with merge fan-in — and it is the only one of the three that survives a noisy host without being loosened into uselessness. Throughput and CPU ratios are reported, not gated.
The gate
GF_INGEST_FLOOR_GATE=1 cargo bench -p graphforge-storage --bench m6_storage_iomeasures the sweep once and exits non-zero on any breach. It runs inm6-walltimebefore the measured pass, so a breach fails the nightly fast, and writesingest-floor-gate.jsonas a 30-day artifact so the ratchet can be argued from recorded numbers.Deliberately not the epic's 1,000,000 edges/sec. A gate set at a value the code cannot meet is switched off within a week and then protects nothing. Every limit carries a comment saying it ratchets up in the same pull request that wins the gain.
Proof it fails closed. Floor raised to 100,000 edges/sec, nothing else changed:
Restored, and the gate passes again.
Inventory
scripts/ci/check-m6-benchmarks.pypins both new names and now also refuses:missing M6 benchmarks: m6_storage_io.rs:ingest_throughputm6-walltimejob does not run the gate —m6-walltime must run the ingest floor gateingest floor gate is missing limits: INGEST_CEILING_BYTES_READ_PER_EDGEingest_throughput must sweep at least two dataset sizesEach verified by making the change and observing exit 1.
What is not measured
The serialized fraction of the ingest path, the epic's ≤2% target. It is a property of where time is spent inside the path rather than of the path's cost, so it needs a profiler or explicit in-path instrumentation; neither divan nor the session's evidence can supply it. Effective core utilisation is reported as the nearest observable proxy and is explicitly not gated as if it were the measurement. CPU per edge and core utilisation are both captured, from
getrusage, and report as unavailable rather than zero on platforms without it.The CI storage inventory required updating
scripts/ci/test-ci-storage-policy.pykeeps a hand-maintained inventory of every artifact upload name, path and retention across all workflows, so a new upload requires a deliberate entry rather than appearing silently. The gate's measurement upload had none, and the Repository Policy job failed the first CI round because of it — the inventory catching drift at the moment it was introduced, which is what it is for.Registering the name also surfaced two other ways the upload was off-policy, fixed in the same push rather than one CI round at a time:
if-no-files-foundmust beerror, notwarn.The gate registry needs nothing.
config/gate-registry.jsonis a one-to-one inventory of workflow files, andcodspeed.ymlis already registered as thecodspeedgate (classscheduled_health_stress, ownerperformance). The floor gate is a step inside that workflow, not a new one, andmake gate-registry-checkpasses unchanged.Verification
cargo fmt --all -- --check,cargo clippy --workspace -- -D warnings,ruff check/ruff format --check,make workflow-lint,make gate-registry-check, and these Repository Policy suites all pass locally:test-ci-storage-policy.py,test-codspeed-nightly.py,test-require-gates.sh,test-verify-ci-gate-enforcement.py,test-classify-changes.sh,test-classify-policy-suites.sh,source_size_policy.py, andcheck-m6-benchmarks.py.Files touched:
crates/graphforge-storage/benches/m6_storage_io.rs,.github/workflows/codspeed.yml,scripts/ci/check-m6-benchmarks.py,scripts/ci/test-ci-storage-policy.py. Nothing undersrc/, and noCargo.tomlchange — the benchmarks live in the existing walltime bench target rather than a new one.Part of #1387
🤖 Generated with Claude Code
Need help on this PR? Tag
@codesmith-botwith what you need. Autofix is disabled.