Skip to content

feat(storage): publish file-backed graph generations beyond the snapshot envelope #338

Description

@DecisionNerd

Problem

The supported public project path cannot represent the large canonical graphs already exercised by the M4 scale probe. Graph publication currently captures every stable graph-workspace file into memory, stores all file contents in one Arrow BinaryArray, limits each captured file to 1 GiB, and limits the complete graph snapshot to 2 GiB.

The measured 3.1M-node/117M-edge and 8M-node/128M-edge projects occupy approximately 14.5–15.5 GiB. They can be built and queried through lower-level Rust storage/execution paths, but they cannot be committed and reopened through GraphForge::new. Raising constants would retain whole-graph memory amplification and Arrow contiguous-buffer limits rather than fixing the public persistence architecture.

Objective

Publish and reopen graph data as an immutable, file-backed project-generation capability so public embedded GraphForge can represent large canonical Parquet/CSR projects without assembling or hydrating the whole graph as one in-memory Arrow participant.

Debt / regime

  • Debt type: architecture and data
  • Quality regime: A compute

Requirements

  1. Define a versioned graph-generation contract whose canonical inventory records normalized contained file paths, lengths, digests, logical roles, and required schema/version metadata while graph files remain file-backed.
  2. Stage and publish the complete graph capability under the existing immutable generation and atomic CURRENT authority. The graph inventory and every referenced file must be validated before promotion; partial or unverified content must never become current.
  3. Open read-only graph execution directly from the pinned immutable generation or an equivalently bounded file-backed view. GraphForge::new must not read the complete graph into one Vec<u8>, one Arrow binary array, or a mandatory full temporary copy.
  4. Writes must continue to operate on private staged state and publish atomically. Preserve writer contention, idempotency, recovery, rollback/checkpoint, generation leases, stale-parent detection, and corruption behavior.
  5. Preserve compatibility with existing graph snapshot generations or provide an explicit deterministic migration path. Unsupported versions must fail with the established structured capability/version errors.
  6. Ensure project portability, checkpointing, inspection, and cleanup either support the new graph capability in a bounded manner or return an explicit documented structured limitation; they must not silently omit files or reconstruct a partial graph.
  7. Keep paths contained and link-safe. Graph/source data remains project data outside code Git, and no network authority or external object store is required.
  8. Use test(performance): establish the M4 embedded baseline and entry gate #334 evidence and the measured 8M-node/128M-edge fixture to prove the supported public facade, not only lower-level storage APIs.

Acceptance Criteria

  • A versioned file-backed graph-generation contract and canonical inventory are documented and mechanically validated.
  • The measured 8M-node/128M-edge class can be committed, closed, reopened through GraphForge::new, and queried through the public Rust facade under the declared local resource stops.
  • Read-only open performs no whole-graph byte assembly and no mandatory complete graph copy; structural evidence records files/bytes validated, copied, mapped, or opened.
  • Atomic publication, pinned-generation reads, recovery, rollback/checkpoint behavior, writer contention, and stale/corrupt generation failures retain deterministic tests.
  • Existing v1 graph snapshot projects remain readable or have an explicit tested migration with structured unsupported-version behavior.
  • Public Python and Node reopen/query acceptance continues to call the same Rust-owned generation path without binding-side persistence logic.
  • Scale-limit and project-format documentation distinguishes the old 1/2 GiB snapshot envelope from the accepted file-backed evidence without claiming a universal graph-size ceiling.

BDD Completion Scenarios

Scenario: A large committed graph reopens through the public facade

Given a canonical graph larger than the legacy 2 GiB snapshot envelope
When it is published, closed, and reopened with GraphForge::new
Then the pinned generation validates and queries its file-backed graph capability
And open does not assemble or hydrate the complete graph into one in-memory payload.

Scenario: Publication remains atomic across failure

Given one valid current generation and a staged file-backed replacement
When staging, validation, or promotion fails after writing one or more graph files
Then CURRENT still resolves to the prior complete generation
And recovery rejects or cleans the incomplete attempt without exposing a partial graph.

Scenario: Legacy and unsupported generations fail predictably

Given a supported legacy graph snapshot or an unknown future graph capability version
When GraphForge resolves the generation
Then the supported legacy project opens or migrates deterministically
And the unknown version fails with the documented structured capability error.

Implementation Notes

Likely surfaces include crates/graphforge-api/src/graph_snapshot.rs, facade workspace hydration/open, crates/graphforge-storage/src/project_publication.rs, project_generation.rs, project portability/checkpoints, generation recovery tests, and project-format architecture documentation.

Prefer extending publication to accept verified file-backed participant sources or a graph-owned directory capability over placing large graph bytes in a larger Arrow value. Preserve CURRENT as the sole committed-generation authority.

Observability

Record aggregate generation identity, inventory counts, total declared/validated bytes, validation/copy strategy, open duration, and structured failure phase. Do not log local absolute paths, UUID rows, properties, graph contents, or file payloads.

Security And Privacy

Reject absolute paths, traversal, duplicate/non-canonical paths, links, special files, digest/length mismatch, and inventory escape. No network service or authorization surface is added. Cleanup must remain generation-scoped and must never follow unverified paths.

Testing

  • Unit-test inventory canonicalization, path containment, duplicate detection, digest/length verification, version handling, and bounded readers.
  • Integration-test stage/validate/promote/open, exact generation pinning, writer conflict, failpoints, recovery, checkpoint/rollback, cleanup, and corruption.
  • Run a small deterministic CI fixture through public Rust plus thin Python/Node reopen paths.
  • Map the 8M/128M scenario to reproducible manual/scheduled test(performance): establish the M4 embedded baseline and entry gate #334 evidence with commit, hardware, RSS, storage, and result fingerprint.
  • Run formatting, targeted clippy/tests, and repository-required exact-head CI for every changed surface.

Documentation

Update project-format architecture, atomic publication/recovery, public open semantics, portability/checkpoint behavior, benchmark methodology, and scale limits after evidence is accepted.

Non-Goals

  • Distributed storage, remote object stores, a daemon, or multi-node execution.
  • Redesigning Parquet query execution or adjacency CSR construction in this issue.
  • Raising the legacy 1/2 GiB constants while retaining whole-snapshot capture.
  • Storing benchmark graph data or generated project files in the code repository.
  • Claiming a universal graph-size maximum from the 8M/128M fixture.

Related Issues

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    coreCore source code changesenhancementNew feature or request

    Type

    No type

    Projects

    No projects

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions