You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
{{ message }}
Repository navigation
test(performance): establish the M4 embedded baseline and entry gate #334
GraphForge has deterministic traversal and release benchmarks, but it does not yet have one accepted entry baseline that separates query-pipeline parallelism, graph-kernel parallelism, memory amplification, and accelerator overhead. Beginning optimization work without that baseline would make speedup claims hardware-dependent and could hide regressions in ordering, cancellation, bounded LIMIT execution, or embedded memory use.
The current public facade has one supported runtime configuration: a fixed two-worker Tokio runtime, while DataFusion independently uses its existing default partition policy. GraphForge does not yet expose a resource policy capable of executing 1, 2, 4, 8, and automatic thread configurations. The entry gate must record that current implementation honestly without fabricating unsupported configurations or implementing the optimization it is meant to measure.
Objective
Establish the finite, reproducible performance and scale entry contract that every M4 implementation issue uses for before/after evidence. Capture the current fixed-two-worker baseline now, define the future thread-parity matrix, and delegate execution of that matrix to #337 when the bounded resource policy exists. This issue measures and defines the gate; it does not implement performance optimizations.
Debt / regime
Debt type: test/proof and observability
Quality regime: A compute
Requirements
Exercise the public Rust facade for representative execution classes: fixed-hop LIMIT, full scan/count, aggregate/top-N, at least one iterative graph algorithm, exact vector search, and one embedding workload.
Record the currently supported fixed two-worker facade configuration, including the DataFusion partitioning and runtime behavior actually observed. Do not imply that a requested thread count was honored when no public policy can express it.
Define a versioned deferred parity matrix for thread configurations 1, 2, 4, 8, and automatic. feat(api): define a bounded embedded execution resource policy #337 owns making those configurations executable and proving parity after its resource policy lands; unavailable machine configurations must be recorded rather than fabricated.
Capture dataset identity, graph node/edge counts and density, GraphForge commit, build profile, OS, CPU, logical parallelism, memory, and accelerator identity when applicable.
Record wall time as hardware-specific evidence and capture deterministic structural metrics where available: rows read/decoded/materialized, partitions, iterations, output rows, spill bytes, peak RSS, and result fingerprint.
Reuse the existing deterministic synthetic fixtures and current 1M/10M-edge and LiveJournal benchmark paths where applicable. Large external fixtures remain explicit opt-in inputs; CI must not download them implicitly.
Define a short required CI entry matrix and a larger reproducible manual/scheduled matrix. Timing observations alone must not make required CI flaky.
Prove the current supported configuration returns the canonical Arrow schema, ordering, result fingerprint, structured errors, cancellation outcome, and resource-limit behavior. Define the identical evidence contract that feat(api): define a bounded embedded execution resource policy #337 must execute across 1/2/4/8/automatic.
Publish the baseline, reproduction commands, metric definitions, known limitations, and acceptance thresholds that later M4 issues must meet.
Classify the 8M-node/128M-edge local scale report as discovery evidence because it uses lower-level Rust construction/execution and cannot traverse the current public snapshot path; do not present it as an accepted public-facade baseline.
Acceptance Criteria
One versioned benchmark contract names every required workload, fixture, current configuration, deferred configuration, metric, and evidence artifact.
The short entry matrix runs deterministically through the public Rust facade in CI under the current fixed two-worker implementation without timing sleeps, retries, ignored correctness assertions, or weakened limits.
The large matrix can be reproduced locally or manually from documented commands and records hardware plus dataset identity.
Baseline evidence exists for every required workload and distinguishes structural gates, hardware-specific timing observations, and unsupported/deferred thread configurations.
The resulting gate is linked from M4 implementation issues as their before/after evidence source.
BDD Completion Scenarios
Scenario: A maintainer captures the current M4 entry baseline
Given a clean release build and a declared benchmark fixture
When the M4 entry command runs through the public Rust facade
Then it emits the workload result, structural counters, peak-memory evidence, result fingerprint, hardware identity, exact reproduction command, and observed fixed-two-worker configuration
And it does not claim cross-machine timing comparability or unsupported thread-count execution.
Scenario: The future parity matrix is finite and executable by its owner
Given the current facade cannot express 1, 2, 4, 8, or automatic resource policies
When the entry contract defines those configurations
Then each configuration has the exact schema/order/fingerprint/error/cancellation/limit evidence required from #337
And this issue records them as deferred rather than fabricating results or adding a test-only optimization seam.
Scenario: Required CI stays deterministic
Given a noisy shared CI runner
When the short performance entry gate runs
Then correctness and structural bounds determine pass or fail
And wall-clock measurements are reported without becoming flaky absolute thresholds.
Scenario: Large lower-level evidence remains qualified
Given the local 8M-node/128M-edge report cannot use the current committed public snapshot path
When its results are referenced by M4
Then they are labeled lower-level discovery evidence with their hardware, SHA, resource stops, and workload class
And #338 remains responsible for proving that scale through GraphForge::new.
Implementation Notes
Likely surfaces include the existing traversal/release benchmark harnesses, graphforge-api public-facade tests, query-scoped I/O counters, algorithm invocation fingerprints, and release load-matrix reporting. Prefer extending shared evidence formats over creating disconnected per-algorithm scripts.
The baseline must expose the current implementation honestly, including the fixed two-worker facade runtime, eager Parquet materialization, single-partition custom plans, serial algorithm kernels, and CSR-to-map conversion where those remain observable bottlenecks. It may detect and report runtime/partition configuration, but production resource configuration belongs to #337.
Observability
The harness may record aggregate performance counters and hardware metadata only. It must not retain query parameters, properties, UUIDs, graph contents, local project paths, or other user data.
Security And Privacy
No network service or authorization surface is added. External datasets are caller-supplied opt-in fixtures and must not be downloaded or uploaded implicitly. Evidence artifacts must exclude graph content and sensitive local paths.
Testing
Unit-test evidence-schema validation, result-fingerprint stability, and honest classification of current, unavailable, and deferred configurations.
Integration-test the current fixed-two-worker short matrix through the public Rust facade.
Preserve existing traversal demand/cancellation and concurrency gates.
Run formatting, targeted clippy/tests, and repository gates appropriate to every changed surface.
Documentation
Document the benchmark contract, metric meanings, reproduction commands, current fixed runtime, deferred parity ownership, hardware caveats, qualified lower-level scale report, and accepted entry results. Update performance or scale-limit documentation only with evidence produced or explicitly classified by this gate.
Non-Goals
Implementing configurable runtime workers, DataFusion resource policy, streaming scans, CSR representation changes, CPU parallel kernels, or GPU kernels.
Declaring a universal node/edge maximum from one machine.
Making large external datasets mandatory for pull-request CI.
Weakening deterministic results or existing correctness tests to obtain better timing.
Optional External Scale Track (non-blocking)
Graph500 × GSI progressive scale and the LDBC full suite are an optional external evidence track for later M4/post-M4 scale validation. They must not delay this entry gate.
Execution (generators, drivers, first-fail / progressive reporting) lives in an external scale harness, not GraphForge core CI. WDC Hyperlink Graphs are not the scale track (#399–#407 closed as not planned). Full LDBC audit completion does not block this gate or #335.
Problem
GraphForge has deterministic traversal and release benchmarks, but it does not yet have one accepted entry baseline that separates query-pipeline parallelism, graph-kernel parallelism, memory amplification, and accelerator overhead. Beginning optimization work without that baseline would make speedup claims hardware-dependent and could hide regressions in ordering, cancellation, bounded
LIMITexecution, or embedded memory use.The current public facade has one supported runtime configuration: a fixed two-worker Tokio runtime, while DataFusion independently uses its existing default partition policy. GraphForge does not yet expose a resource policy capable of executing
1,2,4,8, and automatic thread configurations. The entry gate must record that current implementation honestly without fabricating unsupported configurations or implementing the optimization it is meant to measure.Objective
Establish the finite, reproducible performance and scale entry contract that every M4 implementation issue uses for before/after evidence. Capture the current fixed-two-worker baseline now, define the future thread-parity matrix, and delegate execution of that matrix to #337 when the bounded resource policy exists. This issue measures and defines the gate; it does not implement performance optimizations.
Debt / regime
Requirements
LIMIT, full scan/count, aggregate/top-N, at least one iterative graph algorithm, exact vector search, and one embedding workload.1,2,4,8, and automatic. feat(api): define a bounded embedded execution resource policy #337 owns making those configurations executable and proving parity after its resource policy lands; unavailable machine configurations must be recorded rather than fabricated.1/2/4/8/automatic.Acceptance Criteria
1/2/4/8/automatic matrix and its exact parity assertions are linked to feat(api): define a bounded embedded execution resource policy #337; closing this entry issue does not claim that unsupported configurations already ran.BDD Completion Scenarios
Scenario: A maintainer captures the current M4 entry baseline
Given a clean release build and a declared benchmark fixture
When the M4 entry command runs through the public Rust facade
Then it emits the workload result, structural counters, peak-memory evidence, result fingerprint, hardware identity, exact reproduction command, and observed fixed-two-worker configuration
And it does not claim cross-machine timing comparability or unsupported thread-count execution.
Scenario: The future parity matrix is finite and executable by its owner
Given the current facade cannot express
1,2,4,8, or automatic resource policiesWhen the entry contract defines those configurations
Then each configuration has the exact schema/order/fingerprint/error/cancellation/limit evidence required from #337
And this issue records them as deferred rather than fabricating results or adding a test-only optimization seam.
Scenario: Required CI stays deterministic
Given a noisy shared CI runner
When the short performance entry gate runs
Then correctness and structural bounds determine pass or fail
And wall-clock measurements are reported without becoming flaky absolute thresholds.
Scenario: Large lower-level evidence remains qualified
Given the local 8M-node/128M-edge report cannot use the current committed public snapshot path
When its results are referenced by M4
Then they are labeled lower-level discovery evidence with their hardware, SHA, resource stops, and workload class
And #338 remains responsible for proving that scale through
GraphForge::new.Implementation Notes
Likely surfaces include the existing traversal/release benchmark harnesses,
graphforge-apipublic-facade tests, query-scoped I/O counters, algorithm invocation fingerprints, and release load-matrix reporting. Prefer extending shared evidence formats over creating disconnected per-algorithm scripts.The baseline must expose the current implementation honestly, including the fixed two-worker facade runtime, eager Parquet materialization, single-partition custom plans, serial algorithm kernels, and CSR-to-map conversion where those remain observable bottlenecks. It may detect and report runtime/partition configuration, but production resource configuration belongs to #337.
Observability
The harness may record aggregate performance counters and hardware metadata only. It must not retain query parameters, properties, UUIDs, graph contents, local project paths, or other user data.
Security And Privacy
No network service or authorization surface is added. External datasets are caller-supplied opt-in fixtures and must not be downloaded or uploaded implicitly. Evidence artifacts must exclude graph content and sensitive local paths.
Testing
Documentation
Document the benchmark contract, metric meanings, reproduction commands, current fixed runtime, deferred parity ownership, hardware caveats, qualified lower-level scale report, and accepted entry results. Update performance or scale-limit documentation only with evidence produced or explicitly classified by this gate.
Non-Goals
1/2/4/8/automatic configurations ran before feat(api): define a bounded embedded execution resource policy #337 makes them supported.Optional External Scale Track (non-blocking)
Graph500 × GSI progressive scale and the LDBC full suite are an optional external evidence track for later M4/post-M4 scale validation. They must not delay this entry gate.
Execution (generators, drivers, first-fail / progressive reporting) lives in an external scale harness, not GraphForge core CI. WDC Hyperlink Graphs are not the scale track (#399–#407 closed as not planned). Full LDBC audit completion does not block this gate or #335.
Related Issues