You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
{{ message }}
Repository navigation
perf(storage): add CodSpeed durability and transaction regression baselines #782
M6 changes GraphForge's durable open, recovery, delta replay, compaction, and transaction paths, but the continuous CodSpeed lane currently measures only graphforge-core canonical primitives and graphforge-cypher compilation. A correctness-green M6 tree could therefore acquire material CPU, allocation, open, commit, replay, or compaction regressions without a versioned comparison against a stable baseline.
Durable I/O also cannot be measured honestly with the existing CPU-simulation instrument alone. Simulation is appropriate for deterministic pure-Rust kernels; file flushes, namespace publication, process coordination, spill, and multithreaded behavior require CodSpeed walltime on a stable macro runner or explicitly separate hardware-bound evidence.
Objective
Add bounded, deterministic CodSpeed coverage for M6's pure storage and transaction kernels, add correctly instrumented macro evidence for durable I/O, and attach exact-head performance results to final M6 certification without turning benchmarks into correctness or merge authority.
Debt / regime
Debt type: test/proof, observability, and infrastructure/toil.
Quality regime: A — deterministic benchmark definitions, provenance, and result interpretation.
Requirements
Add a Divan-compatible graphforge-storage benchmark target and, where it measures distinct public-facade overhead, a graphforge-api transaction target.
Run CodSpeed CPU simulation over deterministic pure kernels: GFDR encode/decode/checksum/replay, manifest and reachability calculation, deterministic delta merge/fingerprint work, and transaction staging/classification.
Measure at declared ladders such as 1/100/10,000 operations or runs; construct fixtures outside the measured closure and use no network, sleeps, wall-clock-dependent input, or user data.
Measure open/recovery, durable commit, garbage collection, spill, and compaction with CodSpeed walltime on a runner intended for stable walltime measurement. Do not interpret those I/O paths through CPU simulation.
Add CodSpeed memory instrumentation for replay/compaction allocation and peak-memory evidence when the supported runner is available; otherwise retain an explicit scheduled hardware-bound memory artifact and document the limitation.
Freeze a pre-M6 baseline commit and compare the final exact integrated M6 head against it. Record benchmark names, modes, fixture versions, result URLs/artifacts, and any accepted regression rationale.
graphforge-storage M6 benchmarks run in CodSpeed simulation on pull requests to main and pushes to main.
Pure benchmark groups cover journal framing/verification/replay, reachability, delta merge/fingerprint, and transaction staging/classification across declared size ladders.
Durable open/commit/recovery/GC/compaction evidence uses CodSpeed walltime on a suitable stable runner and identifies the exact host/instrument contract.
Replay and compaction have allocation/peak-memory evidence through CodSpeed memory mode or an explicitly documented scheduled fallback when that mode is unavailable.
The benchmark target inventory, Bazel exception ledger, dependency fingerprint, locks, workflow documentation, and local commands remain coherent.
No benchmark result is presented as durability, atomicity, boundedness, or isolation proof.
BDD Completion Scenarios
Given a pull request changes journal or replay code, when CodSpeed runs, then the deterministic CPU comparison reports the affected size ladders against the base commit.
Given an M6 path performs file flushes, namespace publication, spill, or process coordination, when performance is measured, then walltime runs on the declared stable runner rather than being inferred from CPU simulation.
Given the final integrated M6 commit, when certification closes, then exact-head CPU, walltime, and memory evidence is linked with every material regression repaired or explicitly dispositioned.
Given a benchmark fails or regresses, when it is triaged, then no retry, removed sample, weakened fixture, or inflated threshold converts the signal into false success.
Implementation Notes
Likely surfaces include crates/graphforge-storage/benches/, an optional crates/graphforge-api/benches/ target, crate manifests, .github/workflows/codspeed.yml, Makefile, docs/development/benchmarking.md, and the Bazel migration/exception inventory governed by #729.
Prefer small versioned fixtures that exercise the same Rust-owned primitives used by #752/#753. End-to-end durable I/O cases should use private temporary roots and keep setup, teardown, and validation outside the measured region when the harness supports it.
Observability
CodSpeed result URLs, benchmark identity, measurement mode, base/head commits, fixture version, size ladder, and accepted-regression disposition are the performance evidence. Never include graph contents, query/property values, credentials, or sensitive host paths.
Security And Privacy
Use synthetic content and private contained roots. Preserve OIDC authentication; do not add long-lived CodSpeed credentials. Benchmarks must not follow links, inspect unrelated filesystem state, or upload datasets.
Testing
Prove benchmark discovery locally through ordinary Divan execution and cargo codspeed build/run.
Validate workflow syntax, Cargo/Bazel drift, the benchmark target ledger, and deterministic fixture/version contracts.
Keep correctness assertions outside timed closures and map benchmark fixtures back to the relevant deterministic storage/API tests.
Verify exact benchmark count and names in CI so accidental loss of M6 coverage fails closed.
Documentation
Update the canonical benchmarking guide with the M6 benchmark groups, simulation/walltime/memory decision, runner contract, local commands, diagnostic-only status, baseline selection, and regression-disposition procedure.
Non-Goals
Making CodSpeed a required correctness or merge gate.
A database leaderboard, universal hardware claim, or billion-edge capacity certification.
Problem
M6 changes GraphForge's durable open, recovery, delta replay, compaction, and transaction paths, but the continuous CodSpeed lane currently measures only
graphforge-corecanonical primitives andgraphforge-cyphercompilation. A correctness-green M6 tree could therefore acquire material CPU, allocation, open, commit, replay, or compaction regressions without a versioned comparison against a stable baseline.Durable I/O also cannot be measured honestly with the existing CPU-simulation instrument alone. Simulation is appropriate for deterministic pure-Rust kernels; file flushes, namespace publication, process coordination, spill, and multithreaded behavior require CodSpeed walltime on a stable macro runner or explicitly separate hardware-bound evidence.
Objective
Add bounded, deterministic CodSpeed coverage for M6's pure storage and transaction kernels, add correctly instrumented macro evidence for durable I/O, and attach exact-head performance results to final M6 certification without turning benchmarks into correctness or merge authority.
Debt / regime
Requirements
graphforge-storagebenchmark target and, where it measures distinct public-facade overhead, agraphforge-apitransaction target.//:ci_rust_tests, deterministic fault tests, and native platform tests remain correctness authority.Acceptance Criteria
graphforge-storageM6 benchmarks run in CodSpeed simulation on pull requests tomainand pushes tomain.BDD Completion Scenarios
Implementation Notes
Likely surfaces include
crates/graphforge-storage/benches/, an optionalcrates/graphforge-api/benches/target, crate manifests,.github/workflows/codspeed.yml,Makefile,docs/development/benchmarking.md, and the Bazel migration/exception inventory governed by #729.Prefer small versioned fixtures that exercise the same Rust-owned primitives used by #752/#753. End-to-end durable I/O cases should use private temporary roots and keep setup, teardown, and validation outside the measured region when the harness supports it.
Observability
CodSpeed result URLs, benchmark identity, measurement mode, base/head commits, fixture version, size ladder, and accepted-regression disposition are the performance evidence. Never include graph contents, query/property values, credentials, or sensitive host paths.
Security And Privacy
Use synthetic content and private contained roots. Preserve OIDC authentication; do not add long-lived CodSpeed credentials. Benchmarks must not follow links, inspect unrelated filesystem state, or upload datasets.
Testing
cargo codspeed build/run.Documentation
Update the canonical benchmarking guide with the M6 benchmark groups, simulation/walltime/memory decision, runner contract, local commands, diagnostic-only status, baseline selection, and regression-disposition procedure.
Non-Goals
Related Issues
Open Questions
None.