Skip to content

fix(storage): S22 construction fails "control record exceeds bound" because the shape-end checkpoint ledger holds one entry per sealed segment (#1519) #1526

Description

@DecisionNerd

Symptom

Current main 4fbfe84c fails the G500 host ladder at S22 in the ingest phase, deterministically. ab1a713e (two commits earlier) passes S18–S22 on the same host, same inputs, same quiet-host window (2026-09-21).

gf --json --project <p> import-session validate --session-uuid 00000000-0000-4000-8000-000000000022
{"error":{"code":"GF_IO","message":"storage error: graph construction session: control record exceeds bound","details":{"source":"runtime","kind":"storage"}}}
exit 3, after ~295 s

S18, S19, S20 pass on 4fbfe84c. The ladder evidence is /home/ubuntu/graphforge-ladder/clean-4fbfe84c-evidence/ (with s22-failure-raw/), the abandoned work root is /home/ubuntu/graphforge-ladder/clean-4fbfe84c-work/workspace/s22/, and a hand replay with stderr is /home/ubuntu/graphforge-ladder/s22-ingest-repro-4fbfe84c/ (run.sh, stderr.txt, staging dir intact).

Cause

shape.rs finishes the shape by writing shape-intent.json (bounded by MAX_SHAPE_CONTROL_BYTES = 32 MiB, 8.9 MB at S22, succeeds) and then replace_checkpoint_control (shape.rs:956), which is bounded by MAX_CONTROL_BYTES = 1 MiB. That call clears storage_allocation_transitions but keeps storage_active_identity_allocated_bytes, and at S22 that map is now:

entries serialized bytes
checkpoint before shaping 4,288 265,856
evidence at shape end (from shape-intent.json.final_evidence) 50,128 3,057,304
MAX_CONTROL_BYTES 1,048,576

The 50,128 entries are one per sealed partition segment: the staging dir holds 50,115 part-* segment files (37,632 endpoints, 4,306 identities, 4,002 edge-details, 3,856 resolved, 319 node-details) across 210 shape-progress-* boundaries. #1519 (4fbfe84c) changed sealed: Vec<Option<ArtifactReceipt>> to Vec<Vec<ArtifactReceipt>> and seals every open spill at each boundary, so the segment count now grows with routed input (boundaries × partitions) instead of being bounded by partition_count (4,096). Before #1519 the same map was at most ~4 families × 4,096 partitions ≈ 16.6k entries ≈ 1.0 MB — already within a few percent of the bound at S22, which is why the old tree passed and S22 is the first rung to fail now.

#1519 was measured at S17/S18 only (its commit message), where there are too few boundaries to reach the bound.

Acceptance criteria

  • S22 ingest on the fixed tree passes validate and commit with the same result digests as ab1a713e (the S22 rung on /home/ubuntu/graphforge-ladder/clean-ab1a713e-evidence/ is the reference).
  • The durable checkpoint's size no longer grows with the number of sealed segments (either the per-segment allocation entries are reconciled/rolled up at the boundary that retires them, or they are not part of the checkpoint), so the bound holds at S26 by construction, not by margin.
  • A regression test constructs a shape with enough boundaries × partitions to exceed 1 MiB of per-segment entries under the old behaviour and asserts the shape-end checkpoint write succeeds and stays under MAX_CONTROL_BYTES.
  • Not acceptable: raising MAX_CONTROL_BYTES alone. The checkpoint is rewritten on every accepted chunk; its size must stay independent of input volume (the comment on replace_checkpoint_control already states this).

Related

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions