Skip to content

epic(performance): close M4 embedded performance and scale #335

Description

@DecisionNerd

Reopened after closure audit

The 2026-08-14 close is superseded. #345's ledger froze commit 2b1bfda8145c7f9fe125a75faf695792fc6bbfcb before the later #498/#499–#588 M4 expansion, while #336 and #338 still lack their stated large-scale outcome evidence. This tracker remains open until #336 and #338 produce that evidence or receive explicit documented maintainer waivers, and #345 reconciles the actual post-#498 final tree.

M5 is tracked separately by #735 and milestone 5. Its >=1-billion-live-edge and hardened-interchange work does not waive, replace, or retroactively satisfy M4.

Tracker Purpose

Coordinate and close M4: Embedded Performance and Scale for v0.5.x as one finite, evidence-led workstream. GraphForge must increase practical scale and throughput without requiring a server, abandoning the Rust-owned engine, or weakening deterministic correctness.

This is the canonical milestone close gate. Every other M4 issue is a native sub-issue and directly blocks this issue. Implementation order follows the live blocked-by relationships summarized below; issue numbers alone do not authorize work out of order.

Scope And Decision

This plan covers the complete accepted non-GPU performance strategy:

  • honest entry and exit evidence;
  • file-backed public graph generations;
  • bounded adjacency construction and streaming Parquet execution;
  • CSR-native graph execution and bounded Arrow result shaping;
  • one embedded resource policy;
  • deterministic CPU parallelism for exact cosine KNN, PageRank, and Node2Vec walk generation;
  • evidence-backed scale, architecture, API, benchmark, and contributor documentation.

GPU acceleration is intentionally excluded from this finite issue ledger. The milestone description permits an optional spike, but M4 closure does not require one. Any future GPU work requires separate research, explicit approval, its own capability/parity contract, and new tracking rather than being inferred from this plan.

Entry Gate

Detailed Execution Plan

The Direct blockers column matches the live GitHub blocked-by graph. Transitive dependencies are not substituted for direct relationships.

Phase Issue Bounded outcome Direct blockers
0 — Entry #334 Versioned public-facade baseline, fixed-two-worker evidence, structural counters, deferred thread matrix None
1 — Policy #337 One per-instance policy for runtime workers, DataFusion partitions, memory/spill, I/O, heavy-query admission, and future local CPU pools #334
1 — Public persistence #338 File-backed immutable graph generations beyond the legacy 1 GiB/file and 2 GiB/snapshot envelope #334
2 — Bounded index build #336 Stream/spill adjacency construction beyond the 134,217,727-row Arrow boundary #334, #337
2 — Relational execution #339 Partitioned Parquet execution with execution-time I/O, bounded memory/spill, pushdown, and preserved fixed-hop demand #334, #337
2 — Columnar output #341 Direct bounded Arrow shaping for every analyst verb without complete row-oriented duplication #334, #339
3 — Native graph representation #340 Traversal and analyst projections consume persisted CSR/dense columns without O(E) hash-map expansion #334, #336, #337
4 — Independent-source CPU #342 Exact/all-score/filtered cosine KNN parallel by source with serial dot-product order and exact merge #334, #337, #340, #341
4 — Iterative CPU #343 PageRank destination work parallelized with canonical contribution/reduction order and identical convergence #334, #337, #340, #341
4 — Deterministic generation CPU #344 Node2Vec walk corpus parallelized by stable start/walk ordinal while training remains serial #334, #337, #340, #341
5 — Exit #345 Final-tree public-facade rerun, combined evidence ledger, documentation reconciliation, and live issue-graph audit #334, #336–#344
6 — Milestone close #335 Close only after every child, including #345, is closed and the final evidence reconciliation is accepted #334, #336–#345

Dependency Flow

The diagram shows the critical implementation flow. In addition, every M4 child directly blocks #335, and #345 is directly blocked by every earlier child as listed in the table.

flowchart TD
    E["#334 Entry baseline"]
    R["#337 Embedded resource policy"]
    G["#338 File-backed graph generations"]
    A["#336 Bounded adjacency build"]
    P["#339 Streaming Parquet execution"]
    O["#341 Bounded Arrow shaping"]
    C["#340 CSR-native execution"]
    K["#342 Exact cosine KNN"]
    PR["#343 Deterministic PageRank"]
    N["#344 Node2Vec walk generation"]
    X["#345 Final exit evidence"]
    T["#335 M4 close gate"]

    E --> R
    E --> G
    E --> A
    E --> P
    E --> O
    E --> C
    E --> K
    E --> PR
    E --> N
    E --> X

    R --> A
    R --> P
    R --> C
    A --> C
    P --> O

    R --> K
    R --> PR
    R --> N
    C --> K
    C --> PR
    C --> N
    O --> K
    O --> PR
    O --> N

    G --> X
    A --> X
    P --> X
    O --> X
    C --> X
    K --> X
    PR --> X
    N --> X
    R --> X

    X --> T
Loading

Workstream Contracts

Capacity before compute

#338, #336, #339, #341, and #340 remove public-persistence, contiguous-array, eager-I/O, row-duplication, and graph-representation ceilings. More threads must not be used to conceal those structural limits.

One resource authority

#337 is the sole owner of instance compute/memory/spill/concurrency policy. #336 and #339 consume it for bounded storage/execution; #342–#344 consume its private local CPU pool. No issue may introduce an uncoordinated global worker pool or separate spill authority.

Kernel-specific deterministic parallelism

CPU acceleration is divided by semantics rather than a blanket parallel-iterator conversion:

Each retains a measured serial crossover path for small workloads.

Final evidence is a real gate

#345 performs no optimization. It freezes the accepted tree, reruns #334 through the public facade, verifies relationship and CI evidence, corrects documentation, and is the only child authorized to hand the completed ledger to #335.

Maintenance Rules

  • Native sub-issues and live blocked-by edges are the authoritative membership and execution ledger; keep them consistent with the table and Mermaid flow.
  • Attach every new M4 issue as a native child of epic(performance): close M4 embedded performance and scale #335 and make it directly block epic(performance): close M4 embedded performance and scale #335 before work begins.
  • New work belongs in M4 only when it satisfies an existing acceptance criterion or isolates a verified blocker. It must remain bounded, avoid overlap, and block the canonical close path.
  • Do not create per-log-line issues, CI-only busywork, speculative algorithm slices, or GPU work under this plan.
  • Preserve unrelated branches, worktrees, files, builds, and issue ownership. Sequence implementation from live dependencies.
  • Update this body only when scope or dependency contracts change; routine progress is represented by native issue state and relationships.

Closure Standard

Problem

Performance work can appear successful through selected microbenchmarks while leaving public snapshot ceilings, whole-file materialization, peak-memory amplification, deterministic-order changes, binding regressions, or small-graph latency regressions unresolved. M4 needs one explicit close gate that reconciles all children against the same embedded-first outcome.

Objective

Close M4 only after the repository proves a meaningfully larger or faster embedded operating envelope on the accepted workload classes, with bounded resources and unchanged public correctness contracts.

Requirements

  1. Every M4 implementation issue is a native child, blocks this tracker, and is closed through merged work or an explicit evidence-backed non-code disposition.
  2. test(performance): reconcile final M4 exit evidence and scale claims #345 reruns the test(performance): establish the M4 embedded baseline and entry gate #334 contract as exit evidence on the final accepted tree for every affected workload class.
  3. CPU-only embedded GraphForge remains the universal complete runtime with no required daemon, network service, GPU, Python algorithm engine, or Node algorithm engine.
  4. Rust continues to own persistence, execution, algorithms, resource policy, and Arrow semantics; Python and Node remain thin adapters.
  5. Thread selection preserves canonical schemas, row ordering, deterministic values/fingerprints, cancellation, resource limits, structured errors, atomic publication, and recovery behavior.
  6. Performance claims distinguish deterministic structural evidence, hardware-specific timings, and measured crossover thresholds.
  7. Documentation reports only scale and speedups reproduced by the accepted gate and names the hardware, dataset, graph layout, resource policy, and workload class.

Acceptance Criteria

BDD Completion Scenarios

Scenario: CPU-only embedded operation remains complete

Given a supported machine with no GraphForge service or accelerator
When a caller opens a project and executes Cypher or an analyst verb
Then GraphForge runs in process through the Rust engine and returns canonical Arrow results
And no network authority or foreign graph engine is required.

Scenario: Parallel execution preserves the public contract

Given one committed graph, invocation, and resource policy
When the same supported operation runs at each accepted thread count
Then schema, ordering, deterministic result fingerprint, errors, cancellation, and limits are identical
And only performance/resource observations may differ.

Scenario: Larger work remains publicly representable and bounded

Given a workload above the prior public snapshot or in-memory envelope and an explicit resource policy
When an in-scope publication, index build, streaming query, or analyst invocation runs
Then it completes or fails with the documented structured resource outcome without legacy whole-project eager materialization
And peak memory, spill, decoded work, and public reopen evidence are recorded.

Scenario: M4 closes with finite evidence

Given every M4 child is closed and the final tree is frozen for reconciliation
When #345 evaluates the entry harness, correctness surfaces, issue graph, and documentation
Then #335 links complete before/after evidence and has no unresolved milestone issue, dependency, regression, or undocumented waiver.

Observability

Performance evidence may contain aggregate counters, build identity, hardware identity, dataset taxonomy ID, normalized resource policy, and result fingerprints. It must not retain graph contents, query parameters, UUIDs, properties, local paths, or sensitive system details beyond what reproduction requires.

Security And Privacy

M4 adds no network authority. Resource and spill configuration must fail closed against unsafe paths, links, traversal, or unbounded storage. Diagnostics and evidence must not expose user graph data.

Testing

Documentation

Update project format/open behavior, storage and execution architecture, resource configuration, algorithm determinism, performance methodology, scale limits, public API/binding guidance, and contributor instructions. Claims must identify dataset and hardware context.

Non-Goals

  • GPU acceleration in this finite M4 issue plan.
  • Distributed or multi-node execution.
  • Turning GraphForge Core into a server or remote authority.
  • Python, Node, NetworkX, igraph, cuGraph, or another external graph engine as a runtime fallback.
  • Parallelizing every algorithm irrespective of semantics and measured benefit.
  • Claiming a universal maximum graph size, service-level objective, or cross-machine timing guarantee.

Related Issues

Optional Graph500 / LDBC Validation Track (does not block close)

External Graph500 × GSI and LDBC full-suite work is filed for discoverability under M4 but is not part of the finite close ledger above:

Issue Role
#408 Graph500 × GSI size-ladder specification
#409 LDBC full-suite benchmark specification
#410 External Graph500 + LDBC harness contract

Do not add these as close blockers for #335 unless a future milestone plan explicitly widens M4. Full LDBC audit completion does not block M4. WDC Hyperlink Graph issues (#399–#407) are closed as not planned.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    enhancementNew feature or request

    Type

    No type

    Projects

    No projects

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions