Skip to content

test(ci): measure Rust binding adapters through native acceptance suites #359

Description

@DecisionNerd

Problem

GraphForge's workspace Rust coverage currently reports 86.46% line coverage, but the aggregate hides two very different realities:

  • Rust crates excluding the Python and Node binding crates measure 93.19% (154,934 / 166,257 covered lines).
  • graphforge-bindings-py measures 0.61% (31 / 5,067).
  • graphforge-bindings-node measures 10.48% (551 / 5,258).

The Python and Node acceptance suites execute real same-SHA native artifacts, but scripts/coverage-rust.sh only collects profiles from Cargo's Rust test processes. Most PyO3 and napi-rs adapter execution is therefore invisible to the Rust coverage ledger. The 85% aggregate can pass while thousands of binding-adapter lines appear uncovered, and it cannot distinguish a genuine adapter test gap from a measurement gap.

This is test/proof and CI tooling debt. Coverage must describe executed behavior accurately before it is used to prioritize more tests or raise thresholds.

Objective

Make Rust coverage include the real Python and Node native acceptance paths and publish an honest, deterministic per-surface ledger for core Rust, Python adapter Rust, and Node adapter Rust.

Debt / Regime

  • Debt type: test/proof and infrastructure/toil
  • Quality regime: A — deterministic compute

Requirements

  • Extend the existing cargo llvm-cov harness rather than introducing a second coverage framework.
  • Build same-SHA PyO3 and napi-rs artifacts with LLVM coverage instrumentation in isolated native build output.
  • Run the existing functional Python and Node acceptance suites against those exact instrumented artifacts, including behavioral, persistence/reopen, lifecycle-error, parity, and no-fallback coverage.
  • Merge the native-process profiles into one reproducible report without dropping files, hiding missed lines, or counting a binding-side fallback as Rust execution.
  • Report machine-readable and human-readable totals separately for:
    • Rust crates excluding binding adapters;
    • graphforge-bindings-py Rust code;
    • graphforge-bindings-node Rust code;
    • the complete workspace.
  • Add deterministic validation that each required surface contributed non-empty profile data from the expected same-SHA artifact.
  • Establish an initial 80% line-coverage floor for each Rust binding adapter, based on real acceptance execution. Any verified unreachable/generated exception must be narrow, documented, mechanically checked, and excluded from neither behavioral acceptance nor public-surface parity.
  • Preserve the existing Rust-authoritative architecture: Python and Node remain conversion-only adapters and never become fallback engines.
  • Use isolated build output and never run more than two heavy Rust builds concurrently.
  • Keep local and CI commands practical; reuse already-built instrumented artifacts within a coverage run rather than rebuilding per binding suite.

Acceptance Criteria

  • Python acceptance tests execute an instrumented same-SHA PyO3 artifact and contribute non-zero coverage to graphforge-bindings-py.
  • Node acceptance tests execute an instrumented same-SHA napi-rs artifact and contribute non-zero coverage to graphforge-bindings-node.
  • A machine-readable ledger reports core, Python-adapter, Node-adapter, and whole-workspace totals with source SHA and toolchain identity.
  • The harness fails closed when either adapter profile is missing, stale, empty, or produced by a different artifact/SHA.
  • Both adapter Rust surfaces meet at least 80% line coverage through functional acceptance behavior.
  • Existing Rust, Python, and Node public behavior, parity policies, structured errors, and persistence/reopen suites remain green.
  • make coverage-rust, make check-coverage-rust, and make pre-push use the documented ledger consistently.
  • Coverage documentation explains what each total includes and prevents the whole-workspace number from masking a failed surface.

BDD Completion Scenarios

Scenario: Python behavior contributes Rust adapter coverage

Given a same-SHA instrumented PyO3 artifact and the functional Python acceptance suite
When Rust coverage runs
Then executed Python adapter functions contribute profile data to graphforge-bindings-py
And the suite still asserts exact Rust-owned results, reopen behavior, and structured errors.

Scenario: Node behavior contributes Rust adapter coverage

Given a same-SHA instrumented napi-rs artifact and the functional Node acceptance suite
When Rust coverage runs
Then executed Node adapter functions contribute profile data to graphforge-bindings-node
And the suite still rejects fallback, missing, pending, or manufactured results.

Scenario: Missing binding evidence fails closed

Given a coverage run where one adapter artifact or profile is absent, stale, or empty
When the coverage ledger is validated
Then the run fails with a surface-specific diagnostic
And the core Rust percentage cannot conceal the missing binding evidence.

Scenario: Adapter regression breaches its own floor

Given complete core coverage and one adapter below its accepted 80% floor
When the coverage gate runs
Then the gate fails for that adapter
And no unrelated surface can average the failure away.

Implementation Notes

Likely affected surfaces include:

  • scripts/coverage-rust.sh
  • scripts/check-coverage-rust.sh
  • Makefile
  • Python maturin build/acceptance commands
  • Node napi build/acceptance commands
  • .github/workflows/test.yml if hosted coverage enforcement changes
  • docs/engineering/TESTING.md

Use LLVM's existing profile environment and report-merging facilities. Validate actual artifact identity rather than assuming a package import resolved to the intended build.

Observability

No product telemetry is required. Coverage artifacts may record aggregate counts, source SHA, toolchain version, artifact identity, and command outcome. They must not record graph contents, properties, query text, local user paths, secrets, or package credentials.

Security And Privacy

No new runtime or network surface is expected. Instrumented artifacts must use existing local fixtures and must not be published. Coverage profiles and logs must remain secret-free and data-minimized.

Testing

  • Unit/mutation tests for missing, empty, stale, and wrong-surface profile evidence.
  • Real Python binding behavioral and persistence/reopen suites against the instrumented artifact.
  • Real Node native/BDD behavioral and persistence/reopen suites against the instrumented artifact.
  • Deterministic parser tests for per-surface coverage floors and malformed summaries.
  • Targeted iteration followed by cargo fmt --all -- --check, cargo clippy --workspace -- -D warnings, cargo test --workspace, and make pre-push.

Documentation

Update docs/engineering/TESTING.md and Make target help so contributors understand the core, adapter, and aggregate totals, prerequisites, outputs, and failure diagnostics.

Non-Goals

  • Adding product behavior or binding-side implementations.
  • Excluding binding adapters merely to raise the workspace percentage.
  • Replacing cargo llvm-cov with a new coverage framework.
  • Raising core-engine coverage; that is tracked separately after this ledger is trustworthy.
  • Treating line coverage alone as proof of correctness.

Related Issues

Open Questions

None.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    ci-cdCI/CD configuration changesrelease:noneNo release note or version impacttestingTest coverage and testing infrastructuretoolingDeveloper tooling and automation

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions