Skip to content

perf(bench): measure ingest throughput continuously and gate on a floor - #1406

Merged
DecisionNerd merged 2 commits into
mainfrom
feat/1387-ingest-throughput-bench
Sep 17, 2026
Merged

DecisionNerd merged 2 commits into
mainfrom
feat/1387-ingest-throughput-bench

Conversation

@DecisionNerd

@DecisionNerd DecisionNerd commented Sep 17, 2026 •

Copy link
Copy Markdown
Contributor

Workstream 6 of #1387. Ingest had no throughput target anywhere in the project and no continuous regression detection. Every gain the redesign wins without this is unprotected.

What lands

Two divan benchmarks in the walltime file, on the isolated CodSpeed Macro Runner alongside the other durable I/O cases, because ingest does real durability work that CPU simulation cannot see. Both stage and publish one generation through the same GraphConstructionSession path the scale ladder's ingest phase uses, so the numbers are directly comparable to the ladder's.

  • ingest_throughput — 524,288 and 8,388,608 edges at a fixed edge factor of 16.
  • ingest_identifier_density — 131,072 edges at fan-out 2 and 256, the structural axis.

Plus a fail-closed floor gate in the same binary, run by the nightly, and both names pinned in the benchmark inventory.

Sizes, and why

A sixteen-fold span, the same shape as the ladder's S18-to-S22 fall.

The lower rung is not smaller because below about half a million edges the fixed cost of opening and publishing a generation dominates and the ratio stops describing scaling: across 131,072 to 2,097,152 edges the read-amplification ratio is 1.046, while across 524,288 to 8,388,608 it is 1.128. The upper rung is capped by the nightly budget, not by what would be most informative — the ladder's 67.1M rung takes about fifteen minutes. All four ingest benchmarks together cost 1m58s, and the gate's own pass costs about three minutes.

Cost of the upper rung: 3.6 GiB of transient construction disk, reported by the gate so the runner requirement stays visible. The job timeout moves 60 → 90 minutes.

Measured

Gate output, this branch, on a 16-core x86 development host:

     edges   vertices    wall_s   edges/sec  read_B/edge  write_B/edge  cpu_us/edge   cores    peak_MiB
    524288      32768    12.359       42421         2139           782        10.30    0.44       201.3
   8388608     524288   180.576       46455         2412          1043        12.66    0.59      3604.7
16x size span: bytes read per edge 1.128x (limit 1.20x), throughput 0.913x, cpu per edge 1.229x
ingest floor gate: pass

On an unloaded host the divan pass published 8,388,608 edges in 1.799 min, 77,672 edges/sec, against the ladder's 72,989 at 67.1M and 79,431 at 8.4M. CPU per edge lands at 10.30 to 12.66 µs against the epic's 11.93. The benchmark is measuring the path the epic measured.

The structural axis earns its place: at a constant 131,072 edges, fan-out 256 publishes in 1.2 s and fan-out 2 in 1.885 s, a 1.57x spread from identifier-space density alone that a pure size sweep would have blurred.

The ratio is gated on bytes read per edge

The requirement is to track the ratio between sizes. It is evaluated on bytes read per edge rather than on edges per second, because across four runs of the same commit:

524,288 8,388,608 ratio
bytes read per edge 2138.937 every run 2412.016 every run 1.128 every run
CPU µs per edge 9.84 – 10.57 12.66 – 13.82 1.229 – 1.307
edges/sec 23,285 – 65,155 30,644 – 46,455 0.913 – 1.480

Bytes read per edge reproduced to the byte while wall-clock throughput moved by 48% on identical code. It expresses the same defect — the ladder's throughput fall is read amplification growing with merge fan-in — and it is the only one of the three that survives a noisy host without being loosened into uselessness. Throughput and CPU ratios are reported, not gated.

The gate

GF_INGEST_FLOOR_GATE=1 cargo bench -p graphforge-storage --bench m6_storage_io measures the sweep once and exits non-zero on any breach. It runs in m6-walltime before the measured pass, so a breach fails the nightly fast, and writes ingest-floor-gate.json as a 30-day artifact so the ratchet can be argued from recorded numbers.

Limit Value Measured Note
throughput floor, every size 15,000 edges/s 23,285 – 65,155 provisional and coarse; wall clock is the only metric a busy host can depress without a code change
bytes read per edge, every size 3,000 2,139 / 2,412 deterministic; #1384 targets an order of magnitude off it
CPU µs per edge, every size 15.0 9.84 – 13.82 just under the 15.14 measured before the last two storage improvements, so that 21% cannot be given back silently
bytes read per edge, small → large 1.20x 1.128x deterministic

Deliberately not the epic's 1,000,000 edges/sec. A gate set at a value the code cannot meet is switched off within a week and then protects nothing. Every limit carries a comment saying it ratchets up in the same pull request that wins the gain.

Proof it fails closed. Floor raised to 100,000 edges/sec, nothing else changed:

ingest floor gate: 524288 edges: 36471 edges/sec is below the 100000 edges/sec floor
ingest floor gate: 8388608 edges: 37833 edges/sec is below the 100000 edges/sec floor
EXIT=1

Restored, and the gate passes again.

Inventory

scripts/ci/check-m6-benchmarks.py pins both new names and now also refuses:

  • a bench function deleted from the source — missing M6 benchmarks: m6_storage_io.rs:ingest_throughput
  • a workflow whose m6-walltime job does not run the gate — m6-walltime must run the ingest floor gate
  • a gate missing any of its four limits — ingest floor gate is missing limits: INGEST_CEILING_BYTES_READ_PER_EDGE
  • a sweep collapsed to one size — ingest_throughput must sweep at least two dataset sizes

Each verified by making the change and observing exit 1.

What is not measured

The serialized fraction of the ingest path, the epic's ≤2% target. It is a property of where time is spent inside the path rather than of the path's cost, so it needs a profiler or explicit in-path instrumentation; neither divan nor the session's evidence can supply it. Effective core utilisation is reported as the nearest observable proxy and is explicitly not gated as if it were the measurement. CPU per edge and core utilisation are both captured, from getrusage, and report as unavailable rather than zero on platforms without it.

The CI storage inventory required updating

scripts/ci/test-ci-storage-policy.py keeps a hand-maintained inventory of every artifact upload name, path and retention across all workflows, so a new upload requires a deliberate entry rather than appearing silently. The gate's measurement upload had none, and the Repository Policy job failed the first CI round because of it — the inventory catching drift at the moment it was introduced, which is what it is for.

Registering the name also surfaced two other ways the upload was off-policy, fixed in the same push rather than one CI round at a time:

  • if-no-files-found must be error, not warn.
  • retention is one day for a transfer artifact, not thirty. The gate's measurements are neither a publication nor a certification artifact; their durable record is the CodSpeed walltime series and the gate step's own log.

The gate registry needs nothing. config/gate-registry.json is a one-to-one inventory of workflow files, and codspeed.yml is already registered as the codspeed gate (class scheduled_health_stress, owner performance). The floor gate is a step inside that workflow, not a new one, and make gate-registry-check passes unchanged.

Verification

cargo fmt --all -- --check, cargo clippy --workspace -- -D warnings, ruff check / ruff format --check, make workflow-lint, make gate-registry-check, and these Repository Policy suites all pass locally: test-ci-storage-policy.py, test-codspeed-nightly.py, test-require-gates.sh, test-verify-ci-gate-enforcement.py, test-classify-changes.sh, test-classify-policy-suites.sh, source_size_policy.py, and check-m6-benchmarks.py.

Files touched: crates/graphforge-storage/benches/m6_storage_io.rs, .github/workflows/codspeed.yml, scripts/ci/check-m6-benchmarks.py, scripts/ci/test-ci-storage-policy.py. Nothing under src/, and no Cargo.toml change — the benchmarks live in the existing walltime bench target rather than a new one.

Part of #1387

🤖 Generated with Claude Code


View with [code]smith Autofix with [code]smith
Need help on this PR? Tag @codesmith-bot with what you need. Autofix is disabled.

GraphForge had no throughput target anywhere and no continuous regression
detection on ingest. The existing durable cases cover open, commit, recovery
scan, reachability scan, garbage collection and spill compaction; none of them
ingests in bulk. The three storage improvements that made the previous evidence
stale were each validated by a bespoke one-off measurement, which is why
re-running the whole scale ladder was the only way to learn where we stood.

Add `ingest_throughput` and `ingest_identifier_density` to the walltime
benchmark file, on the isolated CodSpeed Macro Runner alongside the other
durable I/O cases, because ingest does real durability work that CPU simulation
cannot see. Both stage and publish one generation through the same
`GraphConstructionSession` path the scale ladder's ingest phase uses, so the
numbers are comparable to the ladder's: 8,388,608 edges publishes at 77,672
edges/sec here against the ladder's 72,989 at 67.1M.

Two sizes, not one. The defect is that throughput degrades as the graph grows,
and a benchmark at a single fixed size would report a flat healthy number while
that continued underneath it. The sweep runs 524,288 and 8,388,608 edges, a
sixteen-fold span chosen because below about half a million edges the fixed cost
of publishing a generation dominates and the ratio stops describing scaling,
while the ladder's 67.1M rung takes fifteen minutes and cannot run nightly. All
four ingest benchmarks together cost two minutes.

Report bytes read per edge as a first-class metric: 2,139 at the small size and
2,412 at the large one, against 265 bytes retained per edge, the same constant
overhead the epic measured as 4,579 at 67.1M. Report CPU microseconds per edge
and effective core utilisation too, from `getrusage`: 10.30 and 12.66 µs/edge,
0.44 and 0.59 of 16 cores. The serialized fraction of the path is not measurable
from divan or the session's evidence and is not faked.

Gate on the floor, not only on regression. CodSpeed compares each run against
the previous one, so a slow drift that never regresses in a single step passes
indefinitely. `GF_INGEST_FLOOR_GATE` turns the bench binary into a fail-closed
gate that runs in the nightly before the measured pass. It fails when throughput
at any size drops below a floor, when bytes read or CPU per edge at any size
exceed a ceiling, or when bytes read per edge grows more than 1.20x across the
sweep. The ratio is evaluated on bytes read per edge because that quantity
reproduced to the byte on four runs while wall-clock throughput moved by 48% on
the same code; the limits are set at today's measured values and each carries a
comment saying it ratchets as the redesign lands.

Register both names in the fail-closed inventory, which now also refuses a
workflow that does not run the gate, a gate missing any of its four limits, and
a sweep collapsed to a single size.

Part of #1387

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
@coderabbitai

coderabbitai Bot commented Sep 17, 2026 •

Copy link
Copy Markdown

Important

Review skipped

Auto reviews are disabled on this repository. Please check the settings in the CodeRabbit UI or the .coderabbit.yaml file in this repository. To trigger a single review, invoke the @coderabbitai review command.

⚙️ Run configuration

Configuration used: Path: .coderabbit.yaml

Review profile: CHILL

Plan: Advanced

Run ID: 0c493393-a109-4a65-a76a-6807badbf40f

You can disable this status message by setting the reviews.review_status to false in the CodeRabbit configuration file.

Use the checkbox below for a quick retry:

  • 🔍 Trigger review

Warning

Billing warning: we have not been able to collect payment for this subscription for more than 72 hours. Please update the payment method or pay any pending invoices in Billing to avoid service interruption.


Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@github-actions github-actions Bot added core Core source code changes testing Test coverage and testing infrastructure documentation Improvements or additions to documentation ci-cd CI/CD configuration changes tooling Developer tooling and automation labels Sep 17, 2026
…ventory

`scripts/ci/test-ci-storage-policy.py` keeps a hand-maintained inventory of
every artifact upload name, path and retention across all workflows, so that a
new one requires a deliberate entry rather than appearing silently. The ingest
floor gate's measurement upload had none, and the Repository Policy job caught
it at the moment it was introduced, which is what the inventory is for.

Register the name and its path, and conform the upload to the policy it was
breaking in two other ways it would have failed on next: `if-no-files-found`
must be `error`, not `warn`, and retention is one day for a transfer artifact.
Thirty days was wrong — the gate's measurements are neither a publication nor a
certification artifact, and their durable record is the CodSpeed walltime series
plus the gate step's own log.

The gate registry needs nothing: `config/gate-registry.json` is a one-to-one
inventory of workflow files and `codspeed.yml` is already registered as the
`codspeed` gate. The floor gate is a step inside it, not a new workflow, and
`make gate-registry-check` passes unchanged.

Part of #1387

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

ci-cd CI/CD configuration changes core Core source code changes documentation Improvements or additions to documentation testing Test coverage and testing infrastructure tooling Developer tooling and automation

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant