Skip to content

test(tck): extract an in-process Divan benchmark over the TCK scenarios #1653

Description

@DecisionNerd

Parent and bounded concern

Native sub-issue of #1467, split per its XL scope to satisfy existing criteria only. This sub-issue supplies the Divan in-process per-scenario measurement boundary that #1467's revised body requires. It contains no threshold logic, baseline, or comparison. Those belong to the sibling sub-issue (BenchExec whole-TCK run plus a provenance-gated consumer). #1467 remains the close gate.

Scope

  • Move the shared TCK harness pieces (GraphForgeWorld, corpus normalization such as copy_features_normalized) out of crates/graphforge-api/tests/bdd/main.rs into shared modules, with no behaviour change to the Cucumber correctness run or passing_baseline.txt.
  • Add an in-process Divan bench target in graphforge-api (harness = false) that parses the normalized corpus and executes each scenario's steps through the same registered step functions (GraphForgeWorld::collection().find(step)). There is no second semantics engine and no benchmark-only product API.
  • Use the same pooled fixture and clear-on-lease semantics as the Cucumber run. The timed region is the whole scenario, including Given an empty graph, so fixture reset cost stays visible.
  • Check each sample's pass/fail verdict outside the timed region. A failing scenario aborts the run and never yields timing.
  • Scenario keys use the existing <feature>:<line>:<name> format.
  • Machine-readable output comes from CodSpeed's walltime raw_results (written under CODSPEED_ENV). Maintainer decision on test(tck): the performance baseline is from a CI runner, so the perf gate warns on every local run #1467: this file is accepted as Divan evidence, because Divan measures and CodSpeed only serializes. Record that clarification in docs/development/benchmarking.md.
  • Register the target in config/benchmark-measurement-inventory.json as framework authority, and update the build metadata for whichever build authority is current (see decision(build): make one build system the CI authority; measure Bazel remote cache against Cargo with sccache #1618's Bazel retirement).

Acceptance criteria

  • The bench runs the full TCK scenario set in-process and emits walltime raw results keyed by scenario. Test-mode execution produces no performance evidence.
  • A deliberately failing step aborts the bench instead of recording timing (direct test).
  • The Cucumber TCK correctness gate is unchanged: 3897/3897 passing and passing_baseline.txt semantics identical.
  • scripts/ci/benchmark-measurement-policy.py passes with the new inventory entry.

Non-goals

Thresholds, baselines, provenance matching, BenchExec wrapping, and outlier triage all belong to #1467 or its sibling sub-issue.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions