Skip to content

fix(storage): stream topology rewrites within a bounded memory window #902

Description

@DecisionNerd

Problem

The Fly S20 OOM in #900 is a GraphForge storage-path defect, not a host or harness artifact. Each bulk edge publication opened the writer by fully decoding topology to recover surrogate tails, then decoded every existing fixed-schema Parquet row and concatenated it with the new batch before rewriting. Anonymous RSS therefore grew with accumulated graph size.

Controlled public-facade diagnostics using fixed 1,048,576-edge publication windows measured:

  • S17 before repair: 963,362,816-byte peak RSS
  • S18 before repair: 1,369,128,960-byte peak RSS
  • the second S17 publication copied all 1,048,576 existing edges
  • S18 emitted 10,358,282 cumulative topology rows for roughly 3.8 million retained edges

The cumulative recopy I/O is a separate architectural defect tracked by #901. This issue isolates the bounded-memory repair required before that append-only redesign.

Objective

Bound writer-open and fixed-schema rewrite memory by the Parquet row-group / publication window rather than accumulated topology size, while preserving ordinary project semantics and legacy node-schema compatibility.

Acceptance criteria

  • Writer reopen recovers node and edge surrogate tails without full topology reads.
  • Fixed-schema node and edge rewrites stream existing Parquet through bounded batches rather than concatenate the complete prior file in memory.
  • Legacy scalar type_id node topology is normalized batch-by-batch while streaming.
  • Failed staging leaves the prior destination untouched and the staged replacement still commits atomically under the existing contract.
  • Deterministic storage tests assert bounded rewrite batch size and zero full-topology reads on reopen.
  • The Graph500 ladder journals ingest subphase, chunk index, anonymous/file RSS, topology-work counters, and disk usage during long ingest phases.
  • Identical lower-rung diagnostics demonstrate RSS plateauing with accumulated edge count; no claim is made that cumulative I/O or retained disk is linear until feat(storage): stage topology append-only with linear ingest I/O #901 closes.

Evidence target

The repaired S18 diagnostic should stay bounded near the configured request/node working set as additional edge windows publish. The same 4 GiB Fly class must complete S20 before #900 can close.

Non-goals

This issue does not implement append-only topology generations, persistent endpoint indexing, or linear aggregate I/O; those are the root storage-construction outcome in #901.

Blocks #900. Related to #745.

Metadata

Metadata

Assignees

No one assigned

    Labels

    bugSomething isn't working

    Type

    No type

    Projects

    No projects

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions