Skip to content

test(tck): extract an in-process Divan benchmark over the TCK scenarios - #1676

Merged
DecisionNerd merged 1 commit into
mainfrom
test/1653-divan-tck-bench
Sep 30, 2026
Merged

DecisionNerd merged 1 commit into
mainfrom
test/1653-divan-tck-bench

Conversation

@DecisionNerd

@DecisionNerd DecisionNerd commented Sep 30, 2026 •

Copy link
Copy Markdown
Contributor

Description

Adds the in-process Divan per-scenario measurement boundary for the openCypher TCK that #1467 requires. There are no thresholds, baselines or comparisons here; those belong to #1654.

Closes #1653
Part of #1467

Changes Made

  • Shared harness modules, no behaviour change. GraphForgeWorld moves to tests/bdd/world.rs. The corpus normalization (copy_features_normalized, normalize_leading_continuations, tck_only_filter) and the <feature>:<line>:<name> key moves to tests/bdd/corpus.rs. tests/bdd/main.rs uses them unchanged.
  • Divan bench benches/tck_scenarios/ (harness = false):
    • It normalizes the same corpus and parses it with Cucumber's own parser, so outline expansion matches the correctness run.
    • It resolves every step through GraphForgeWorld::collection().find(step) and calls the registered step function. Panics are caught the same way Cucumber catches them.
    • It uses the same pooled fixture with clear-on-lease: fixture::activate(), and fixture::release where the Cucumber after hook calls it. The engine count is still asserted <= TCK_CONCURRENCY.
    • The timed region is the whole scenario, including Given an empty graph. Each iteration's verdict is checked outside the timed region, before the next sample and before the bench function returns. Divan writes a benchmark's raw results only after that function returns, so a failing scenario panics first and yields no timing.
    • Benchmarks are named scenario[<feature>:<line>:<name>]. The defaults are sample_count = 10 and sample_size = 1.
  • Machine-readable output is CodSpeed walltime raw_results under CODSPEED_ENV. The maintainer's clarification is recorded in docs/development/benchmarking.md.
  • Inventory: both bench files are registered in config/benchmark-measurement-inventory.json as framework_authority. make bench-tck-scenarios was added. Cargo is the only build description (ADR 0048). The target has no required-features, so there is no gated-target lane change.
  • Direct test tests/tck_scenario_bench.rs (nextest-runnable). It re-runs its own binary as a child that runs the real Divan bench function in-process with CODSPEED_ENV set, then inspects the exit status and the raw results.
    • A passing scenario yields exactly one raw result, keyed by scenario, with rounds == 3 and iter_per_round == 1. This is the known positive.
    • A deliberately failing step (Then the result should be empty after RETURN 1) makes the run fail with the verdict message and no timing for that scenario.
    • Test mode writes no raw results, and still fails on a failing scenario.

.github/workflows/codspeed.yml and config/gate-registry.json are untouched. The PR CI Gate is unchanged.

Testing

See the verification comment for commands and results.

🤖 Generated with Claude Code


View with [code]smith Autofix with [code]smith
Need help on this PR? Tag @codesmith-bot with what you need. Autofix is disabled.

…os (#1653)

Move GraphForgeWorld and the corpus normalization out of the bdd runner
into shared modules, and add a Divan bench target that executes every
TCK scenario through the same registered step functions and pooled
fixture. Each scenario's verdict is checked outside the timed region;
a failing scenario aborts before Divan records timing.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
@coderabbitai

coderabbitai Bot commented Sep 30, 2026

Copy link
Copy Markdown

Important

Review skipped

Auto reviews are disabled on this repository. Please check the settings in the CodeRabbit UI or the .coderabbit.yaml file in this repository. To trigger a single review, invoke the @coderabbitai review command.

⚙️ Run configuration

Configuration used: Repository: CurateLabs/graphforge/.coderabbit.yaml

Review profile: CHILL

Plan: Advanced

Run ID: 85df3fb2-06c6-41df-bc7c-6c1b911f0221

You can disable this status message by setting the reviews.review_status to false in the CodeRabbit configuration file.

Use the checkbox below for a quick retry:

  • 🔍 Trigger review

Warning

Billing warning: we have not been able to collect payment for this subscription for more than 72 hours. Please update the payment method or pay any pending invoices in Billing to avoid service interruption.


Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@github-actions github-actions Bot added core Core source code changes documentation Improvements or additions to documentation labels Sep 30, 2026
@DecisionNerd

Copy link
Copy Markdown
Contributor Author

Local verification at 9857321f5

Host: shared 16-thread bench host, not quiet (load 7 to 12). None of the timings below are performance claims.

TCK correctness gate, before and after the extraction. Command: python3 scripts/test_environment.py -- cargo test --workspace --locked --test bdd

before (edc91bc9e, origin/main) after (9857321f5)
API BDD 118 passed, 0 failed, 0 skipped 118 passed, 0 failed, 0 skipped
openCypher TCK 3898 passing of 3898, baseline 3898 (0 regressed, 0 xpass) 3898 passing of 3898, baseline 3898 (0 regressed, 0 xpass)
fixture profile pooled-isolated-serial-v1, concurrency 1, engines created 1 same
  • BLESS_TCK_BASELINE=1 on the after tree rewrote tests/tck/passing_baseline.txt byte-identically: sha256 2acf833c…88ec, no git diff.
  • The after bdd binary differs from the before binary (sha256 5b994181… vs fd63a532…), and only the after binary references tests/bdd/corpus.rs. So the second run exercised the extracted code.

Bench, test mode: CODSPEED_ENV=local CODSPEED_CARGO_WORKSPACE_ROOT=<scratch> cargo bench -p graphforge-api --locked --bench tck_scenarios -- --test

  • 3898 scenarios executed, exit 0, 0 raw result files.

Bench, measurement mode: the same command without --test.

  • Exit 0 after 9m53s (release build already cached). It wrote 3898 raw results.
  • Names are unique, and the key set equals passing_baseline.txt exactly.
  • Every result has rounds = 10 and iter_per_round = 1.

Sample raw result:

{"name": "scenario[Delete5 - Delete clause interoperation with built-in data types:34:[1] Delete node from a list]",
 "uri": "crates/graphforge-api/benches/tck_scenarios/runner.rs::runner::scenario[Delete5 - Delete clause interoperation with built-in data types:34:[1] Delete node from a list]",
 "stats": {"min_ns": 68150142, "median_ns": 70764628, "max_ns": 77151670, "rounds": 10, "iter_per_round": 1}}

Create3:49:[2] WITH-CREATE had a median of 71.2 ms. This bench uses the release bench profile, while the Cucumber timings use the debug test profile. These numbers are not outlier triage; that stays on #1467.

Direct test, and mutation proof. Command: cargo test -p graphforge-api --locked --test tck_scenario_bench. Result: 4 passed, 1 ignored (the subprocess entry point).

mutation killed by
M1: require_passed never checks the verdict failing_step_aborts… and test_mode… fail
M2: a panicking step is reported as passed failing_step_aborts… and test_mode… fail
M3: test mode runs run_benches test_mode… fails: "test mode wrote performance evidence: [scenario[BenchFault:3:[1] Passing scenario]]"
M4: child without CODSPEED_ENV passing_scenario_emits… fails. This is the known positive. Without it, the two negative tests would pass vacuously.

Other gates:

  • cargo fmt --all -- --check: clean.
  • cargo clippy --workspace -- -D warnings: clean.
  • cargo clippy -p graphforge-api --bench tck_scenarios --test tck_scenario_bench -- -D warnings: no findings in the new or moved files. It does report pre-existing pedantic findings in the included api_steps.rs and tck_steps.rs, which the bdd target compiles too. That check is stricter than CI.
  • make pre-push-fast: passed.
  • benchmark-measurement-policy.py: 20 sites verified. Its tests pass.
  • test-ci-storage-policy.py, docs-tree-policy.py check and make gate-registry-check: OK.

@DecisionNerd

Copy link
Copy Markdown
Contributor Author

Integration note: closingIssuesReferences is empty even though the body says Closes #1653. GitHub is currently not linking closing keywords on any open PR (#1673, #1674, #1677 show the same), so this is platform-side. The intended closing issue is #1653 alone. I'll close it by hand with the squash commit after merge.

Reviewed at 9857321f5:

  • The bench reuses the bdd step modules via #[path], so there is no second step engine.
  • The Cucumber gate shows 3898/3898 before and after, and passing_baseline.txt regenerates byte-identically.
  • The failing-step abort was shown by breaking the code on purpose, and nextest does not run the bench target.
  • CI Gate passed at the exact head, merge state is CLEAN, and there are no open threads.

@DecisionNerd
DecisionNerd added this pull request to the merge queue Sep 30, 2026
Merged via the queue into main with commit bf28be7 Sep 30, 2026
26 checks passed
@DecisionNerd
DecisionNerd deleted the test/1653-divan-tck-bench branch September 30, 2026 13:12
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

core Core source code changes documentation Improvements or additions to documentation

Projects

None yet

Development

Successfully merging this pull request may close these issues.

test(tck): extract an in-process Divan benchmark over the TCK scenarios

1 participant