You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
The supported public project path cannot represent the large canonical graphs already exercised by the M4 scale probe. Graph publication currently captures every stable graph-workspace file into memory, stores all file contents in one Arrow BinaryArray, limits each captured file to 1 GiB, and limits the complete graph snapshot to 2 GiB.
The measured 3.1M-node/117M-edge and 8M-node/128M-edge projects occupy approximately 14.5–15.5 GiB. They can be built and queried through lower-level Rust storage/execution paths, but they cannot be committed and reopened through GraphForge::new. Raising constants would retain whole-graph memory amplification and Arrow contiguous-buffer limits rather than fixing the public persistence architecture.
Objective
Publish and reopen graph data as an immutable, file-backed project-generation capability so public embedded GraphForge can represent large canonical Parquet/CSR projects without assembling or hydrating the whole graph as one in-memory Arrow participant.
Debt / regime
Debt type: architecture and data
Quality regime: A compute
Requirements
Define a versioned graph-generation contract whose canonical inventory records normalized contained file paths, lengths, digests, logical roles, and required schema/version metadata while graph files remain file-backed.
Stage and publish the complete graph capability under the existing immutable generation and atomic CURRENT authority. The graph inventory and every referenced file must be validated before promotion; partial or unverified content must never become current.
Open read-only graph execution directly from the pinned immutable generation or an equivalently bounded file-backed view. GraphForge::new must not read the complete graph into one Vec<u8>, one Arrow binary array, or a mandatory full temporary copy.
Writes must continue to operate on private staged state and publish atomically. Preserve writer contention, idempotency, recovery, rollback/checkpoint, generation leases, stale-parent detection, and corruption behavior.
Preserve compatibility with existing graph snapshot generations or provide an explicit deterministic migration path. Unsupported versions must fail with the established structured capability/version errors.
Ensure project portability, checkpointing, inspection, and cleanup either support the new graph capability in a bounded manner or return an explicit documented structured limitation; they must not silently omit files or reconstruct a partial graph.
Keep paths contained and link-safe. Graph/source data remains project data outside code Git, and no network authority or external object store is required.
A versioned file-backed graph-generation contract and canonical inventory are documented and mechanically validated.
The measured 8M-node/128M-edge class can be committed, closed, reopened through GraphForge::new, and queried through the public Rust facade under the declared local resource stops.
Read-only open performs no whole-graph byte assembly and no mandatory complete graph copy; structural evidence records files/bytes validated, copied, mapped, or opened.
Existing v1 graph snapshot projects remain readable or have an explicit tested migration with structured unsupported-version behavior.
Public Python and Node reopen/query acceptance continues to call the same Rust-owned generation path without binding-side persistence logic.
Scale-limit and project-format documentation distinguishes the old 1/2 GiB snapshot envelope from the accepted file-backed evidence without claiming a universal graph-size ceiling.
BDD Completion Scenarios
Scenario: A large committed graph reopens through the public facade
Given a canonical graph larger than the legacy 2 GiB snapshot envelope
When it is published, closed, and reopened with GraphForge::new
Then the pinned generation validates and queries its file-backed graph capability
And open does not assemble or hydrate the complete graph into one in-memory payload.
Scenario: Publication remains atomic across failure
Given one valid current generation and a staged file-backed replacement
When staging, validation, or promotion fails after writing one or more graph files
Then CURRENT still resolves to the prior complete generation
And recovery rejects or cleans the incomplete attempt without exposing a partial graph.
Scenario: Legacy and unsupported generations fail predictably
Given a supported legacy graph snapshot or an unknown future graph capability version
When GraphForge resolves the generation
Then the supported legacy project opens or migrates deterministically
And the unknown version fails with the documented structured capability error.
Implementation Notes
Likely surfaces include crates/graphforge-api/src/graph_snapshot.rs, facade workspace hydration/open, crates/graphforge-storage/src/project_publication.rs, project_generation.rs, project portability/checkpoints, generation recovery tests, and project-format architecture documentation.
Prefer extending publication to accept verified file-backed participant sources or a graph-owned directory capability over placing large graph bytes in a larger Arrow value. Preserve CURRENT as the sole committed-generation authority.
Observability
Record aggregate generation identity, inventory counts, total declared/validated bytes, validation/copy strategy, open duration, and structured failure phase. Do not log local absolute paths, UUID rows, properties, graph contents, or file payloads.
Security And Privacy
Reject absolute paths, traversal, duplicate/non-canonical paths, links, special files, digest/length mismatch, and inventory escape. No network service or authorization surface is added. Cleanup must remain generation-scoped and must never follow unverified paths.
Testing
Unit-test inventory canonicalization, path containment, duplicate detection, digest/length verification, version handling, and bounded readers.
Run formatting, targeted clippy/tests, and repository-required exact-head CI for every changed surface.
Documentation
Update project-format architecture, atomic publication/recovery, public open semantics, portability/checkpoint behavior, benchmark methodology, and scale limits after evidence is accepted.
Non-Goals
Distributed storage, remote object stores, a daemon, or multi-node execution.
Redesigning Parquet query execution or adjacency CSR construction in this issue.
Raising the legacy 1/2 GiB constants while retaining whole-snapshot capture.
Storing benchmark graph data or generated project files in the code repository.
Claiming a universal graph-size maximum from the 8M/128M fixture.
Problem
The supported public project path cannot represent the large canonical graphs already exercised by the M4 scale probe. Graph publication currently captures every stable graph-workspace file into memory, stores all file contents in one Arrow
BinaryArray, limits each captured file to 1 GiB, and limits the complete graph snapshot to 2 GiB.The measured 3.1M-node/117M-edge and 8M-node/128M-edge projects occupy approximately 14.5–15.5 GiB. They can be built and queried through lower-level Rust storage/execution paths, but they cannot be committed and reopened through
GraphForge::new. Raising constants would retain whole-graph memory amplification and Arrow contiguous-buffer limits rather than fixing the public persistence architecture.Objective
Publish and reopen graph data as an immutable, file-backed project-generation capability so public embedded GraphForge can represent large canonical Parquet/CSR projects without assembling or hydrating the whole graph as one in-memory Arrow participant.
Debt / regime
Requirements
CURRENTauthority. The graph inventory and every referenced file must be validated before promotion; partial or unverified content must never become current.GraphForge::newmust not read the complete graph into oneVec<u8>, one Arrow binary array, or a mandatory full temporary copy.Acceptance Criteria
GraphForge::new, and queried through the public Rust facade under the declared local resource stops.BDD Completion Scenarios
Scenario: A large committed graph reopens through the public facade
Given a canonical graph larger than the legacy 2 GiB snapshot envelope
When it is published, closed, and reopened with
GraphForge::newThen the pinned generation validates and queries its file-backed graph capability
And open does not assemble or hydrate the complete graph into one in-memory payload.
Scenario: Publication remains atomic across failure
Given one valid current generation and a staged file-backed replacement
When staging, validation, or promotion fails after writing one or more graph files
Then
CURRENTstill resolves to the prior complete generationAnd recovery rejects or cleans the incomplete attempt without exposing a partial graph.
Scenario: Legacy and unsupported generations fail predictably
Given a supported legacy graph snapshot or an unknown future graph capability version
When GraphForge resolves the generation
Then the supported legacy project opens or migrates deterministically
And the unknown version fails with the documented structured capability error.
Implementation Notes
Likely surfaces include
crates/graphforge-api/src/graph_snapshot.rs, facade workspace hydration/open,crates/graphforge-storage/src/project_publication.rs,project_generation.rs, project portability/checkpoints, generation recovery tests, and project-format architecture documentation.Prefer extending publication to accept verified file-backed participant sources or a graph-owned directory capability over placing large graph bytes in a larger Arrow value. Preserve
CURRENTas the sole committed-generation authority.Observability
Record aggregate generation identity, inventory counts, total declared/validated bytes, validation/copy strategy, open duration, and structured failure phase. Do not log local absolute paths, UUID rows, properties, graph contents, or file payloads.
Security And Privacy
Reject absolute paths, traversal, duplicate/non-canonical paths, links, special files, digest/length mismatch, and inventory escape. No network service or authorization surface is added. Cleanup must remain generation-scoped and must never follow unverified paths.
Testing
Documentation
Update project-format architecture, atomic publication/recovery, public open semantics, portability/checkpoint behavior, benchmark methodology, and scale limits after evidence is accepted.
Non-Goals
Related Issues