test(tck): extract an in-process Divan benchmark over the TCK scenarios - #1676
Conversation
…os (#1653) Move GraphForgeWorld and the corpus normalization out of the bdd runner into shared modules, and add a Divan bench target that executes every TCK scenario through the same registered step functions and pooled fixture. Each scenario's verdict is checked outside the timed region; a failing scenario aborts before Divan records timing. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
|
Important Review skippedAuto reviews are disabled on this repository. Please check the settings in the CodeRabbit UI or the ⚙️ Run configurationConfiguration used: Repository: CurateLabs/graphforge/.coderabbit.yaml Review profile: CHILL Plan: Advanced Run ID: You can disable this status message by setting the Use the checkbox below for a quick retry:
Warning Billing warning: we have not been able to collect payment for this subscription for more than 72 hours. Please update the payment method or pay any pending invoices in Billing to avoid service interruption. Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out. Comment |
Local verification at
|
before (edc91bc9e, origin/main) |
after (9857321f5) |
|
|---|---|---|
| API BDD | 118 passed, 0 failed, 0 skipped | 118 passed, 0 failed, 0 skipped |
| openCypher TCK | 3898 passing of 3898, baseline 3898 (0 regressed, 0 xpass) | 3898 passing of 3898, baseline 3898 (0 regressed, 0 xpass) |
| fixture profile | pooled-isolated-serial-v1, concurrency 1, engines created 1 | same |
BLESS_TCK_BASELINE=1on the after tree rewrotetests/tck/passing_baseline.txtbyte-identically: sha2562acf833c…88ec, no git diff.- The after
bddbinary differs from the before binary (sha2565b994181…vsfd63a532…), and only the after binary referencestests/bdd/corpus.rs. So the second run exercised the extracted code.
Bench, test mode: CODSPEED_ENV=local CODSPEED_CARGO_WORKSPACE_ROOT=<scratch> cargo bench -p graphforge-api --locked --bench tck_scenarios -- --test
- 3898 scenarios executed, exit 0, 0 raw result files.
Bench, measurement mode: the same command without --test.
- Exit 0 after 9m53s (release build already cached). It wrote 3898 raw results.
- Names are unique, and the key set equals
passing_baseline.txtexactly. - Every result has
rounds = 10anditer_per_round = 1.
Sample raw result:
{"name": "scenario[Delete5 - Delete clause interoperation with built-in data types:34:[1] Delete node from a list]",
"uri": "crates/graphforge-api/benches/tck_scenarios/runner.rs::runner::scenario[Delete5 - Delete clause interoperation with built-in data types:34:[1] Delete node from a list]",
"stats": {"min_ns": 68150142, "median_ns": 70764628, "max_ns": 77151670, "rounds": 10, "iter_per_round": 1}}Create3:49:[2] WITH-CREATE had a median of 71.2 ms. This bench uses the release bench profile, while the Cucumber timings use the debug test profile. These numbers are not outlier triage; that stays on #1467.
Direct test, and mutation proof. Command: cargo test -p graphforge-api --locked --test tck_scenario_bench. Result: 4 passed, 1 ignored (the subprocess entry point).
| mutation | killed by |
|---|---|
M1: require_passed never checks the verdict |
failing_step_aborts… and test_mode… fail |
| M2: a panicking step is reported as passed | failing_step_aborts… and test_mode… fail |
M3: test mode runs run_benches |
test_mode… fails: "test mode wrote performance evidence: [scenario[BenchFault:3:[1] Passing scenario]]" |
M4: child without CODSPEED_ENV |
passing_scenario_emits… fails. This is the known positive. Without it, the two negative tests would pass vacuously. |
Other gates:
cargo fmt --all -- --check: clean.cargo clippy --workspace -- -D warnings: clean.cargo clippy -p graphforge-api --bench tck_scenarios --test tck_scenario_bench -- -D warnings: no findings in the new or moved files. It does report pre-existing pedantic findings in the includedapi_steps.rsandtck_steps.rs, which thebddtarget compiles too. That check is stricter than CI.make pre-push-fast: passed.benchmark-measurement-policy.py: 20 sites verified. Its tests pass.test-ci-storage-policy.py,docs-tree-policy.py checkandmake gate-registry-check: OK.
|
Integration note: Reviewed at
|
Description
Adds the in-process Divan per-scenario measurement boundary for the openCypher TCK that #1467 requires. There are no thresholds, baselines or comparisons here; those belong to #1654.
Closes #1653
Part of #1467
Changes Made
GraphForgeWorldmoves totests/bdd/world.rs. The corpus normalization (copy_features_normalized,normalize_leading_continuations,tck_only_filter) and the<feature>:<line>:<name>key moves totests/bdd/corpus.rs.tests/bdd/main.rsuses them unchanged.benches/tck_scenarios/(harness = false):GraphForgeWorld::collection().find(step)and calls the registered step function. Panics are caught the same way Cucumber catches them.fixture::activate(), andfixture::releasewhere the Cucumberafterhook calls it. The engine count is still asserted<= TCK_CONCURRENCY.Given an empty graph. Each iteration's verdict is checked outside the timed region, before the next sample and before the bench function returns. Divan writes a benchmark's raw results only after that function returns, so a failing scenario panics first and yields no timing.scenario[<feature>:<line>:<name>]. The defaults aresample_count = 10andsample_size = 1.raw_resultsunderCODSPEED_ENV. The maintainer's clarification is recorded indocs/development/benchmarking.md.config/benchmark-measurement-inventory.jsonasframework_authority.make bench-tck-scenarioswas added. Cargo is the only build description (ADR 0048). The target has norequired-features, so there is no gated-target lane change.tests/tck_scenario_bench.rs(nextest-runnable). It re-runs its own binary as a child that runs the real Divan bench function in-process withCODSPEED_ENVset, then inspects the exit status and the raw results.rounds == 3anditer_per_round == 1. This is the known positive.Then the result should be emptyafterRETURN 1) makes the run fail with the verdict message and no timing for that scenario..github/workflows/codspeed.ymlandconfig/gate-registry.jsonare untouched. The PR CI Gate is unchanged.Testing
See the verification comment for commands and results.
🤖 Generated with Claude Code
Need help on this PR? Tag
@codesmith-botwith what you need. Autofix is disabled.