You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
Native sub-issue of #1467, split per its XL scope to satisfy existing criteria only. This sub-issue supplies the Divan in-process per-scenario measurement boundary that #1467's revised body requires. It contains no threshold logic, baseline, or comparison. Those belong to the sibling sub-issue (BenchExec whole-TCK run plus a provenance-gated consumer). #1467 remains the close gate.
Scope
Move the shared TCK harness pieces (GraphForgeWorld, corpus normalization such as copy_features_normalized) out of crates/graphforge-api/tests/bdd/main.rs into shared modules, with no behaviour change to the Cucumber correctness run or passing_baseline.txt.
Add an in-process Divan bench target in graphforge-api (harness = false) that parses the normalized corpus and executes each scenario's steps through the same registered step functions (GraphForgeWorld::collection().find(step)). There is no second semantics engine and no benchmark-only product API.
Use the same pooled fixture and clear-on-lease semantics as the Cucumber run. The timed region is the whole scenario, including Given an empty graph, so fixture reset cost stays visible.
Check each sample's pass/fail verdict outside the timed region. A failing scenario aborts the run and never yields timing.
Scenario keys use the existing <feature>:<line>:<name> format.
The bench runs the full TCK scenario set in-process and emits walltime raw results keyed by scenario. Test-mode execution produces no performance evidence.
A deliberately failing step aborts the bench instead of recording timing (direct test).
The Cucumber TCK correctness gate is unchanged: 3897/3897 passing and passing_baseline.txt semantics identical.
scripts/ci/benchmark-measurement-policy.py passes with the new inventory entry.
Non-goals
Thresholds, baselines, provenance matching, BenchExec wrapping, and outlier triage all belong to #1467 or its sibling sub-issue.
Parent and bounded concern
Native sub-issue of #1467, split per its XL scope to satisfy existing criteria only. This sub-issue supplies the Divan in-process per-scenario measurement boundary that #1467's revised body requires. It contains no threshold logic, baseline, or comparison. Those belong to the sibling sub-issue (BenchExec whole-TCK run plus a provenance-gated consumer). #1467 remains the close gate.
Scope
GraphForgeWorld, corpus normalization such ascopy_features_normalized) out ofcrates/graphforge-api/tests/bdd/main.rsinto shared modules, with no behaviour change to the Cucumber correctness run orpassing_baseline.txt.graphforge-api(harness = false) that parses the normalized corpus and executes each scenario's steps through the same registered step functions (GraphForgeWorld::collection().find(step)). There is no second semantics engine and no benchmark-only product API.Given an empty graph, so fixture reset cost stays visible.<feature>:<line>:<name>format.raw_results(written underCODSPEED_ENV). Maintainer decision on test(tck): the performance baseline is from a CI runner, so the perf gate warns on every local run #1467: this file is accepted as Divan evidence, because Divan measures and CodSpeed only serializes. Record that clarification indocs/development/benchmarking.md.config/benchmark-measurement-inventory.jsonas framework authority, and update the build metadata for whichever build authority is current (see decision(build): make one build system the CI authority; measure Bazel remote cache against Cargo with sccache #1618's Bazel retirement).Acceptance criteria
passing_baseline.txtsemantics identical.scripts/ci/benchmark-measurement-policy.pypasses with the new inventory entry.Non-goals
Thresholds, baselines, provenance matching, BenchExec wrapping, and outlier triage all belong to #1467 or its sibling sub-issue.