Skip to content

test(performance): establish the M4 embedded baseline and entry gate #334

Description

@DecisionNerd

Problem

GraphForge has deterministic traversal and release benchmarks, but it does not yet have one accepted entry baseline that separates query-pipeline parallelism, graph-kernel parallelism, memory amplification, and accelerator overhead. Beginning optimization work without that baseline would make speedup claims hardware-dependent and could hide regressions in ordering, cancellation, bounded LIMIT execution, or embedded memory use.

The current public facade has one supported runtime configuration: a fixed two-worker Tokio runtime, while DataFusion independently uses its existing default partition policy. GraphForge does not yet expose a resource policy capable of executing 1, 2, 4, 8, and automatic thread configurations. The entry gate must record that current implementation honestly without fabricating unsupported configurations or implementing the optimization it is meant to measure.

Objective

Establish the finite, reproducible performance and scale entry contract that every M4 implementation issue uses for before/after evidence. Capture the current fixed-two-worker baseline now, define the future thread-parity matrix, and delegate execution of that matrix to #337 when the bounded resource policy exists. This issue measures and defines the gate; it does not implement performance optimizations.

Debt / regime

  • Debt type: test/proof and observability
  • Quality regime: A compute

Requirements

  1. Exercise the public Rust facade for representative execution classes: fixed-hop LIMIT, full scan/count, aggregate/top-N, at least one iterative graph algorithm, exact vector search, and one embedding workload.
  2. Record the currently supported fixed two-worker facade configuration, including the DataFusion partitioning and runtime behavior actually observed. Do not imply that a requested thread count was honored when no public policy can express it.
  3. Define a versioned deferred parity matrix for thread configurations 1, 2, 4, 8, and automatic. feat(api): define a bounded embedded execution resource policy #337 owns making those configurations executable and proving parity after its resource policy lands; unavailable machine configurations must be recorded rather than fabricated.
  4. Capture dataset identity, graph node/edge counts and density, GraphForge commit, build profile, OS, CPU, logical parallelism, memory, and accelerator identity when applicable.
  5. Record wall time as hardware-specific evidence and capture deterministic structural metrics where available: rows read/decoded/materialized, partitions, iterations, output rows, spill bytes, peak RSS, and result fingerprint.
  6. Reuse the existing deterministic synthetic fixtures and current 1M/10M-edge and LiveJournal benchmark paths where applicable. Large external fixtures remain explicit opt-in inputs; CI must not download them implicitly.
  7. Define a short required CI entry matrix and a larger reproducible manual/scheduled matrix. Timing observations alone must not make required CI flaky.
  8. Prove the current supported configuration returns the canonical Arrow schema, ordering, result fingerprint, structured errors, cancellation outcome, and resource-limit behavior. Define the identical evidence contract that feat(api): define a bounded embedded execution resource policy #337 must execute across 1/2/4/8/automatic.
  9. Publish the baseline, reproduction commands, metric definitions, known limitations, and acceptance thresholds that later M4 issues must meet.
  10. Classify the 8M-node/128M-edge local scale report as discovery evidence because it uses lower-level Rust construction/execution and cannot traverse the current public snapshot path; do not present it as an accepted public-facade baseline.

Acceptance Criteria

  • One versioned benchmark contract names every required workload, fixture, current configuration, deferred configuration, metric, and evidence artifact.
  • The short entry matrix runs deterministically through the public Rust facade in CI under the current fixed two-worker implementation without timing sleeps, retries, ignored correctness assertions, or weakened limits.
  • The large matrix can be reproduced locally or manually from documented commands and records hardware plus dataset identity.
  • Baseline evidence exists for every required workload and distinguishes structural gates, hardware-specific timing observations, and unsupported/deferred thread configurations.
  • The 1/2/4/8/automatic matrix and its exact parity assertions are linked to feat(api): define a bounded embedded execution resource policy #337; closing this entry issue does not claim that unsupported configurations already ran.
  • Peak-memory and spill evidence is captured rather than inferring maximum scale from wall time.
  • The current fixed-hop demand/cancellation contract remains intact and does not acquire eager repartitioning as a side effect of the harness.
  • The lower-level 8M/128M report is retained as explicitly qualified discovery evidence, while public-product claims remain gated by feat(storage): publish file-backed graph generations beyond the snapshot envelope #338 and public-facade reruns.
  • The resulting gate is linked from M4 implementation issues as their before/after evidence source.

BDD Completion Scenarios

Scenario: A maintainer captures the current M4 entry baseline

Given a clean release build and a declared benchmark fixture
When the M4 entry command runs through the public Rust facade
Then it emits the workload result, structural counters, peak-memory evidence, result fingerprint, hardware identity, exact reproduction command, and observed fixed-two-worker configuration
And it does not claim cross-machine timing comparability or unsupported thread-count execution.

Scenario: The future parity matrix is finite and executable by its owner

Given the current facade cannot express 1, 2, 4, 8, or automatic resource policies
When the entry contract defines those configurations
Then each configuration has the exact schema/order/fingerprint/error/cancellation/limit evidence required from #337
And this issue records them as deferred rather than fabricating results or adding a test-only optimization seam.

Scenario: Required CI stays deterministic

Given a noisy shared CI runner
When the short performance entry gate runs
Then correctness and structural bounds determine pass or fail
And wall-clock measurements are reported without becoming flaky absolute thresholds.

Scenario: Large lower-level evidence remains qualified

Given the local 8M-node/128M-edge report cannot use the current committed public snapshot path
When its results are referenced by M4
Then they are labeled lower-level discovery evidence with their hardware, SHA, resource stops, and workload class
And #338 remains responsible for proving that scale through GraphForge::new.

Implementation Notes

Likely surfaces include the existing traversal/release benchmark harnesses, graphforge-api public-facade tests, query-scoped I/O counters, algorithm invocation fingerprints, and release load-matrix reporting. Prefer extending shared evidence formats over creating disconnected per-algorithm scripts.

The baseline must expose the current implementation honestly, including the fixed two-worker facade runtime, eager Parquet materialization, single-partition custom plans, serial algorithm kernels, and CSR-to-map conversion where those remain observable bottlenecks. It may detect and report runtime/partition configuration, but production resource configuration belongs to #337.

Observability

The harness may record aggregate performance counters and hardware metadata only. It must not retain query parameters, properties, UUIDs, graph contents, local project paths, or other user data.

Security And Privacy

No network service or authorization surface is added. External datasets are caller-supplied opt-in fixtures and must not be downloaded or uploaded implicitly. Evidence artifacts must exclude graph content and sensitive local paths.

Testing

  • Unit-test evidence-schema validation, result-fingerprint stability, and honest classification of current, unavailable, and deferred configurations.
  • Integration-test the current fixed-two-worker short matrix through the public Rust facade.
  • Preserve existing traversal demand/cancellation and concurrency gates.
  • Map Scenario 1 to the manual/scheduled large matrix, Scenario 2 to the feat(api): define a bounded embedded execution resource policy #337 parity suite, Scenario 3 to the required short CI job, and Scenario 4 to evidence-schema/documentation validation.
  • Run formatting, targeted clippy/tests, and repository gates appropriate to every changed surface.

Documentation

Document the benchmark contract, metric meanings, reproduction commands, current fixed runtime, deferred parity ownership, hardware caveats, qualified lower-level scale report, and accepted entry results. Update performance or scale-limit documentation only with evidence produced or explicitly classified by this gate.

Non-Goals

  • Implementing configurable runtime workers, DataFusion resource policy, streaming scans, CSR representation changes, CPU parallel kernels, or GPU kernels.
  • Claiming that 1/2/4/8/automatic configurations ran before feat(api): define a bounded embedded execution resource policy #337 makes them supported.
  • Declaring a universal node/edge maximum from one machine.
  • Making large external datasets mandatory for pull-request CI.
  • Weakening deterministic results or existing correctness tests to obtain better timing.

Optional External Scale Track (non-blocking)

Graph500 × GSI progressive scale and the LDBC full suite are an optional external evidence track for later M4/post-M4 scale validation. They must not delay this entry gate.

Execution (generators, drivers, first-fail / progressive reporting) lives in an external scale harness, not GraphForge core CI. WDC Hyperlink Graphs are not the scale track (#399–#407 closed as not planned). Full LDBC audit completion does not block this gate or #335.

Related Issues

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    coreCore source code changesenhancementNew feature or requestexecutorChanges to query executortestingTest coverage and testing infrastructure

    Type

    No type

    Projects

    No projects

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions