perf(storage): range-partition shaping and delete the external merge tree - #1430
Conversation
|
Important Review skippedAuto reviews are disabled on this repository. Please check the settings in the CodeRabbit UI or the ⚙️ Run configurationConfiguration used: Path: .coderabbit.yaml Review profile: CHILL Plan: Advanced Run ID: You can disable this status message by setting the Use the checkbox below for a quick retry:
Warning Billing warning: we have not been able to collect payment for this subscription for more than 72 hours. Please update the payment method or pay any pending invoices in Billing to avoid service interruption. Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out. Comment |
This comment has been minimized.
This comment has been minimized.
1749c85 to
70783c4
Compare
This comment has been minimized.
This comment has been minimized.
7e9f5ed to
83facc9
Compare
…tree Shaping built its global UUID order with a fan-in-32 `BinaryHeap` merge over one run per staged chunk, after copying every staged run once so the merge had an owned input. Both the copies and the tree are gone. Every record is routed once into the partition that owns its key, each partition is sorted independently, and the partitions are concatenated in index order. Because the partition function is monotone, that concatenation *is* the global order: there is nothing left to merge. Two passes replace `1 + ceil(log_32(chunks))`. The partitions run sequentially. No threads are introduced here, deliberately: the determinism contract has to be provable before concurrency exists, so that a later failure can only be concurrency. The splitters are sampled, not computed, and that is settled rather than a preference. `graphforge-core/src/uuid.rs` mints UUIDv7 on every write path, so every identity minted during one ingest shares a near-identical 48-bit timestamp prefix. A formula over the key's high bits is a valid range partition — order is preserved, every digest reproduces — and it routes essentially every row into one partition. Instead, a deterministic systematic sample of the staged identity domain is cut at even quantiles and the chosen splitters are recorded in the shape intent before any partition writes a byte. Sampling is index-driven: the positions are fixed before a byte is read, so the pass seeks to them and costs O(partition_count) records rather than a scan. A collapsed one-partition run is perfectly deterministic and passes every byte-equality test, so determinism tests cannot catch that regression. The partition balance check is what catches it, it refuses at shaping time rather than reporting in a benchmark, and its refusal is proved against real `Uuid::now_v7()` output. Measured: 8192 minted-shape identities at `partition_count = 256` produce 256 effective partitions with mean 32 and max 32 rows, a ratio of 1.00. Formula splitters over 4096 real v7 identities put 4096 of 4096 into one partition and are refused. Peak resident memory on the multilevel allocation workload moves from 108.4-109.0 MB to 147.4-148.9 MB, +36%, bounded by the partition count rather than by data volume. Nothing spills to a bounded per-shard merge: a partition is materialized whole and sorted in memory. Determinism, scoped to what is true. Shaped bytes are not reproducible across sessions on main today, because `runtime_catalog_now_micros` is a wall clock that reaches `shaped-runtime-catalog.parquet` (#1416). Measured on 13632d4 with the merge tree still in place, two sessions over identical input differed in exactly that one artifact and nothing else. So the claim made here is that for a fixed set of recorded session parameters, identical logical input produces byte-identical shaped and encoded artifacts across separate sessions, at every recorded partition count, and across an interrupted-and-resumed run; and that range partitioning contributes no cross-session variance of its own, which `unpinned_sessions_differ_only_in_the_wall_clock_runtime_catalog` asserts directly. The timestamp is not touched here. The partitioned concatenation reproduces the merge tree's fixed-width outputs byte for byte: `shaped-identities.run`, both detail domains and the resolved endpoint domain carry the same digests main produced. Only shaped-rows Parquet moves, and that is the pinned writer properties below. The partition cut is the recorded count bounded by the recorded record count, at a floor of sixteen rows per partition. Each partition costs a durable spill with its own barrier and writer receipt in every family, so cutting the full recorded count regardless of input size made that a constant: a two-thousand-row graph paid a sixty-seven-million-row graph's durability price and shaping's barrier count stopped tracking the work it protects. The ladder measured that directly, 210/278/414 fsyncs on main against 2195/2089/2063. Bounding the cut restores 1143/1579/2063, monotone, with the scale-bearing policy untouched. The floor is the balance check's own minimum, because cutting past it would create partitions the balance assertion cannot validate. The bound is a pure function of the staged record count recorded in the chunk receipts, so it does not make the partition count machine-dependent; at production scale the recorded count is the binding constraint and the bound does nothing. Durable format changes, all intended, no backward compatibility: 1. `GraphConstructionBudgets` gains `partition_count: u32`, default 256, bounded at 4096. It is a recorded format parameter and must never be derived from `available_parallelism()` or a thread count. 2. `ShapeIntent` gains `splitters: Vec<String>` (canonical lower hex, strictly increasing) and `partition_identity_rows: Vec<u64>`. 3. The shaping artifact grammar is replaced. Every `merge-*` family disappears. New durable names: `staged-identities.run`, `staged-endpoints.run`, `shaped-node-details.run`, `shaped-edge-details.run`, `shaped-edge-endpoints.run`, and the per-partition spills `part-<family>-p<NNNNN>.run` and `part-rows-<ns>-p<NNNNN>.arrow`. `ConstructionShape` therefore names `shaped-*` artifacts where it used to name merge roots, which changes the shape inventory and its authority digest. 4. `permanent_parquet::writer_properties()` pins a constant `created_by`. arrow-rs otherwise writes its own crate version into every permanent Parquet footer, putting the library version inside every digest derived from those bytes. This changes every permanent Parquet in the project, not only ingest's. 5. Shaped row Parquet takes those pinned properties. `merge_row_group` passed `None`, so shaped rows used arrow-rs defaults while the runtime catalog did not. That is the hazard the design flagged independently, it was real, and the baseline measurement shows the bytes move when it is fixed. 6. `GraphConstructionEvidence` gains nine partition counters, serde-defaulted. 7. Three shaping failpoints are renamed to their partition equivalents. Unchanged: the encoded graph layout, the CAS object layout, the radix manifest node format, the generation manifest schema, `CURRENT`, and `register-parquet`'s source copy. Part of #1387 Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
83facc9 to
42aaf7e
Compare
… never learned Two defects, both from main moving underneath this branch. `edge_batch` took two more arguments (#1439). The call site passes `node_start: 1` and the fixture's own node count, which reproduces the previous fixed behaviour exactly while `rows <= node_count`, so the measurement is unchanged. The composition grammar predates range-partition shaping (#1430, #1439), so every per-partition spill fell through to `Unclassified` -- 24.4 MB of a 65.5 MB peak at 131k edges and 86.3 MB of 218.0 MB at 524k, the largest single term and precisely the hole the attribution exists to close. Three components now name them: `ShapedPartitionRun` for `part-<family>-p<NNNNN>.run`, `ShapedPartitionRows` for `part-rows-<digest>-p<NNNNN>.arrow`, and `StagedDomainRun` for the whole-domain runs concatenated from those spills. The name grammar is not restated here. Classification calls `partition_shaping::is_partition_artifact_name`, the writer's own predicate, so the two cannot drift apart the way they just did; only the payload split between fixed-width and Arrow is decided at the classification site. `an_unknown_construction_name_is_refused_rather_than_absorbed` still passes, so `Unclassified` is still reachable and this closes the hole rather than hiding it. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
What this does
Shaping built its global UUID order with a fan-in-32
BinaryHeapmerge over one run per staged chunk, after copying every staged run once so the merge had an owned input. Both the copies and the tree are gone. Every record is routed once into the partition that owns its key, each partition is sorted independently, and the partitions are concatenated in index order. Because the partition function is monotone, that concatenation is the global order: there is nothing left to merge. Two passes replace1 + ceil(log_32(chunks)).shaping_merge.rsis deleted (1,672 lines):FixedMergeAccumulator,merge_fixed_group,RowMergeAccumulator,merge_row_group,convert_identity_run,copy_authenticated_runandcopy_authenticated_run_with_codec. All five fixed-width domains and every row-Parquet schema group go through the partitioner.The partitions run sequentially. No threads are introduced, deliberately. The determinism contract has to be provable before concurrency exists, so a later failure can only be concurrency. S6 is the thread step; this is its safety net.
The splitters are sampled, and that is settled
graphforge-core/src/uuid.rsmints UUIDv7 on every write path, so every identity minted during one ingest shares a near-identical 48-bit timestamp prefix. A formula over the key's high bits is a perfectly valid range partition — order is preserved, every digest still reproduces, every byte-equality test still passes — and it routes essentially every row into one partition.So the splitters are data. A deterministic systematic sample of the staged identity domain is sorted, cut at even quantiles, and recorded in the shape intent before any partition writes a byte. Sampling is index-driven: the positions are fixed before a byte is read, so the pass seeks to them and costs
O(partition_count)records rather than a scan of the identity domain.Partition balance is load-bearing, not a nicety
A collapsed one-partition run is perfectly deterministic and passes every byte-equality test in this PR. Only the balance check can tell it apart from a working partition, so it refuses at shaping time, fail-closed, rather than reporting in a benchmark.
Measured, 8192 minted-shape identities at
partition_count = 256:Proved to fail, against real
graphforge_core::uuid::new_v7()output rather than a synthetic distribution —balance_assertion_refuses_a_formula_partitioning_of_minted_identities:The check is scoped to the staged identity domain, which is the domain the splitters are quantiles of. Per-schema row groups and the resolved-endpoint domain are sub-domains, so the quantile guarantee does not transfer to them: on a small mixed fixture one row group measured 5.06x mean while the identity domain measured 1.00x. Asserting on a sub-domain would produce false refusals without catching anything the identity assertion misses.
Determinism, scoped to what is actually true
Shaped bytes are not reproducible across sessions on
maintoday, independent of anything here.ConstructionShape::runtime_catalog_now_microsissession_now_micros, a wall clock, and it reachesshaped-runtime-catalog.parquet, which is a shaped output. That is #1416 and it is not touched here.Baseline measured on unmodified
13632d4b, with the external merge tree still in place — two sessions, identical logical input, ordinary unpinned clock:Pinning
session_now_microsremoved that single difference and nothing else. So the claim this PR makes is precise:unpinned_sessions_differ_only_in_the_wall_clock_runtime_catalogruns two sessions with an ordinary clock and asserts the runtime catalog is the only artifact that differs, so fix(storage): shaped output bytes are not reproducible across sessions, because a wall clock is written into the runtime catalog #1416 cannot be attributed to partitioning and cannot silently grow.Same input twice, identical digests
same_input_twice_produces_identical_digests— two sessions, two project directories, full shape + encode, all 7 shaped and 19 encoded artifacts compared:Two artifacts are excluded, each with a stated reason asserted in the test:
topology/uuid-membership/ordinal-v4-receipt.jsoncarries a freshly minted random rebuild nonce (Uuid::new_v4()). The test asserts it was produced, so the exclusion cannot go stale.shape_authority_sha256serializesArtifactReceipt.identity, which carries volume serial and inode. It has never been reproducible across filesystems. This confirms design §9.4 as fact rather than belief: it is a purely local consistency authority.Different partition counts, identical logical result
partition_count_changes_the_layout_but_not_the_logical_result—[requested P, effective partitions, max identity rows in a partition, identities, peak records materialized at once]:The last row shows the record-count bound at work: 2048 identities admit at most
2048/16 = 128partitions, so a recorded 256 cuts 128. The test asserts that rule rather than a hardcoded list.The layout genuinely differs at every level — partition count, spill set, and the size of the sorted set held in memory falls 16x — and the fingerprint is asserted equal across all four.
This is stronger than the design predicted and it is a deviation worth naming. Design R1 expected changing
Pto change published digests, because §5.4 makes the encoded inventory per-partition. That part is not done here (see below), soPis a staging parameter only and the published bytes do not move. The test asserts byte equality acrossPrather than inequality.Interrupted and resumed
interrupted_shaping_resumes_to_identical_digestscancels shaping mid-pass, reopens, completes, and asserts the full fingerprint equals an uninterrupted run, with identical recorded splitters and partition count.The partitioned output reproduces the merge tree byte for byte
Worth stating on its own:
shaped-identities.run(6ea5b846…), node details (d7eac2fd…), edge details (4b9d9dd1…) and edge endpoints (0b3c8415…) carry the same digestsmainproduced with the external merge tree. The only shaped digests that move are the row Parquets, and that is the writer-properties fix below — which independently confirms that defect was real and load-bearing on the bytes.The
merge_row_grouphazard: confirmed, and worse than flaggedYes, it is the same defect.
merge_row_grouppassedNonefor writer properties (shaping_merge.rs:994), so shaped-rows Parquet used arrow-rs defaults while the runtime catalog usedpermanent_parquet::writer_properties().It is worse than the design states.
writer_properties()never pinnedcreated_byat all, so arrow-rs wrote its own crate version (parquet-rs version X.Y.Z) into the footer of every permanent Parquet in the project — not only ingest's. An arrow-rs bump would have changed the content address of every graph whose logical content did not change, breaking reproducibility and dedup together. Both are fixed here.Durability barriers: a regression found in CI review, and its cause
The first CI run surfaced two things. Both are fixed in this branch.
1. Windows compile failure — not a handle-lifetime defect.
Windows graphforge-storage Locksfailed at its first step witherror[E0433]: cannot find 'unix' in 'os'. Rewritingrecovery/tests.rsonto the replacement publication path dropped a#[cfg(unix)]attribute from the unrelatedsymlink_substitution_is_rejected_on_independent_sealtest that sat between the two tests being replaced. Restored. No production code was involved and no digest moves.An audit of every
#[cfg(...)]in the five test files this PR edits found exactly that one dropped attribute and no others.graphforge-storage --lib --tests,graphforge-apiandgraphforge-filesystemnow all cross-compile clean forx86_64-pc-windows-gnulocally, so the remaining Windows steps that were skipped behind the compile error get a real run this time. Handle ordering was also re-audited end to end: every writer is dropped before its file is renamed or unlinked, including the newDropguards that retire abandoned spills.2. A real durability-barrier regression, caught by
scale_g500_ladder.equivalent_full_lifecycle_1x_2x_4x_has_bounded_metric_policiesfailed withfsync_synchronization.fsync_calls has no positive 1x-to-2x growth: [2195, 2089, 2063]. Measured onmainfor the same rungs: 210, 278, 414 — growing. On the first version of this branch: 2195, 2089, 2063 — flat, and roughly 10x higher at the smallest rung.The cause was structural, not incidental. Each partition costs a durable spill artifact with its own barrier and writer receipt, in every family. Cutting the full recorded 256 partitions regardless of input size made that a constant: a two-thousand-row graph paid the same durability price as a sixty-seven-million-row one, and shaping's barrier count stopped tracking the work it protects.
Fixed at the cause. The cut is now the recorded partition count bounded by the recorded record count, at a floor of 16 rows per partition:
The floor is
BALANCE_MIN_MEAN_ROWS, deliberately: below it the balance check is disabled, so cutting more partitions than that would be creating partitions the balance assertion cannot validate. This does not make the partition count machine-dependent, which is what R1 forbids — the bound is a pure function of the staged record count recorded in the chunk receipts, so the same logical input cuts the same partitions on any host. At production scale the recorded count is the binding constraint and the bound does nothing: 67.1M identities admit 4.1M partitions, so 256 still wins.After the fix: 1143, 1579, 2063 — monotone growth, and the
ScaleBearingpolicy passes untouched. No test assertion was changed, weakened or reclassified.The remaining honest cost: at these rungs the absolute count is still about 5x
main's, because the ladder's 1x-4x rungs sit below the crossover. The merge tree's barrier count grew with chunk count (one install per chunk per family, plus the merge groups); partitioning's is bounded by the partition count. The crossover is at 256 chunks, which is S20 and above — so at every production scale partitioning does strictly fewer barriers, and only sub-256-chunk graphs pay more. Given fsync is 0.25% of ingest at the reference scale and is an explicit non-goal of #1387, this is reported rather than optimised further. If it ever matters, the receipt-per-spill is the lever: spills are transient, are deleted wholesale by incomplete-shape recovery, and are never verified against their receipts within the pass, so their receipts buy a reopen-time check on an artifact reopen discards.Every digest survived the fix. The four fixed-width shaped domains and all 19 encoded artifacts are byte-identical before and after the cut bound — it changes the staged layout only, which is exactly the property the cross-
Ptest already asserts.Two fixtures grew rather than two assertions shrinking
The cut bound made two small fixtures shape into a single partition, so they stopped exercising what they exist for. Relaxing
shape_partitions > 1to>= 1would have made both pass and measured nothing: at one partition there is no range partition, no concatenation across partitions, and no linearity to measure. So the fixtures grew instead:shaping_is_bounded_deterministic_and_multipass_at_1x_2x_4xnow stages 64 nodes and 32 edges per chunk, giving 6 / 12 / 24 partitions across the rungs. It additionally asserts the partition count rises with the input, which is the partitioned analogue of the merge tree gaining a level at 4x — the property the oldmerge_passes >= 2assertion carried.range_partition_work_is_linear_across_production_chunk_countsnow gives each staged chunk 16 records, so the full requested 16 partitions are cut at every chunk count from 31 to 1088, and the exact two-pass work assertions are measured against a real multi-partition concatenation.One partition for a tiny graph is intended, and now asserted
Below 32 identities the cut is 1, and the partitioned path becomes a plain sorted run. That is the correct answer rather than a case to refuse: a partition costs a durable spill with its own barrier and receipt in every family, and refusing the degenerate case would mean refusing small graphs.
a_graph_below_one_partition_of_rows_shapes_into_exactly_one_partitionstates it deliberately — one partition, no splitters, and a fingerprint identical to an explicitlypartition_count = 1run, so the degenerate path is proven to be the same path rather than a second one.Memory
Peak resident memory, same workload and host,
construction_lifecycle_multilevel_allocation_baseline(8,192 nodes + 32,768 edges at 1x/2x/4x), two runs each:13632d4bThe increase is
O(partition_count x buffer) + O(largest partition): 256 open spill writers at a 64 KiB buffer each, plus the one sorted partition materialized at a time. It is bounded by the recorded partition count, not by data volume, so it is a constant, not a multiplier — against the maintainer's 4 GiB budget there is no concern. For comparison, DataFusion paid +52% for the equivalent change.Nothing spills to a bounded per-shard external merge. A partition is materialized whole and sorted in memory (design rule R4), so no write volume returns and #1393's transient-peak projection stands as written.
Wall-clock times are not compared: the host was running three or more concurrent cargo jobs and the baseline binary alone varied 54 s to 99 s between runs.
Durable format changes
Pre-1.0, no backward compatibility. Every one changes on-disk bytes for identical logical input, so every digest changes and older stores are not readable.
GraphConstructionBudgetsgainspartition_count: u32, default 256, validated in1..=4096. It is a recorded format parameter and must never be derived fromavailable_parallelism()or a thread count — that is the single easiest way to break reproducibility across differently-sized hosts.ShapeIntentgainssplitters: Vec<String>(canonical lower hex, strictly increasing) andpartition_identity_rows: Vec<u64>. The splitters are installed before any partition writes a byte, andvalidate_shape_bindingre-validates them as a well-formed monotone range partition every time the intent is read, so a corrupted splitter list fails closed.The shaping artifact grammar is replaced. Every
merge-*family disappears:merge-unified-*,merge-node-source-*,merge-edge-source-*,merge-endpoint-source-*,merge-resolved-source-*,merge-identities-l*-g*,merge-node-details-l*-g*,merge-edge-details-l*-g*,merge-endpoints-l*-g*,merge-resolved-l*-g*,merge-rows-*. New durable names:staged-identities.run,staged-endpoints.run— the pre-surrogate and pre-resolution domains;shaped-node-details.run,shaped-edge-details.run,shaped-edge-endpoints.run;part-<family>-p<NNNNN>.runfor the five fixed families andpart-rows-<ns16>-p<NNNNN>.arrowfor row groups.ConstructionShapetherefore namesshaped-*artifacts where it used to name merge-tree roots, which changes the shape inventory andshape_authority_sha256.permanent_parquet::writer_properties()pins a constantcreated_by,"graphforge permanent parquet/1". This changes the bytes of every permanent Parquet in the project, not only ingest's.Shaped row Parquet takes the pinned properties, where
merge_row_grouppassedNone. Changes shaped-rows bytes with no other change, as the design predicted.GraphConstructionEvidencegains nine partition counters, all#[serde(default)]:shape_partitions,shape_partition_count,splitter_sampled_source_records,splitter_sample_records,max_partition_identity_rows,partitioned_identity_rows,partition_outputs,peak_partition_records,partition_rows.Three shaping failpoints are renamed:
shape.fixed.after_install->shape.partition_spill.after_install,shape.fixed_merge.after_install->shape.partition_output.after_install,shape.row_merge.after_install->shape.row_partition.after_install.Not changed: the encoded graph layout (still one file per logical table), the CAS object layout, the radix manifest node format, the generation manifest schema, the
CURRENTrecord, andregister-parquet's source copy, which stays as deliberate durability policy.Crash, corruption, cancellation and recovery coverage
Unweakened, and two tests are strengthened. Every test whose subject was deleted was ported onto the replacement rather than dropped:
create_unpublished_replaceable_childpublication protocol thatcopy_authenticated_run_with_codecused, soshaping_publication_guard_covers_setup_and_post_rename_failurestransfers verbatim — all ten injection points (initial_file_identity,writer_construction,window_validation,fsync_evidence_overflow,final_file_identity,file_space_usage,install_child,directory_sync,post_publication_metric_overflow,manifest_update) still fire, with the same cleanup-ordering and no-leaked-temp assertions.shaping_copy_failures_finalize_release_and_remove_unpublished_outputsbecomespartition_output_failures_...: same injected input-release and output-cleanup failures, same error ordering, same no-published-artifact assertions.detail_codec_current_format_cancel_corrupt_copy_and_retrybecomes..._cancel_corrupt_route_and_retry. Routing authenticates the staged run against its writer receipt while it streams, so the "source content changed" refusal the copy step provided survives without the copy. Strengthened: it now also asserts that an abandoned partitioner leaves no owned temporary, a newDropguarantee.fixed_merge_work_is_exact_across_production_fan_in_boundariesbecomesrange_partition_work_is_linear_across_production_chunk_counts. Strengthened: the old test pinned logarithmic growth (31 inputs cost 31 records of work, 1025 cost 3073); the new one pins exactly two passes at every size, which is a tighter invariant, across the same 31/32/33/272/1023/1024/1025/1088 boundaries including the S20 and S22 chunk counts.retained_row_roots_cancellation_recovers_and_corruption_fails_closednow damages apart-rows-*.arrowspill; all four damage modes (none / payload / replacement / extra-link) unchanged.shaping_recovery_refuses_same_inode_payload_corruption(fix(storage): authenticate retained shaping bytes during recovery #1269) is untouched and passes both arms.shape.before/after_identity_retirementandshape.before/after_endpoint_retirementare unchanged, and endpoint retirement still happens only after every routed record is durable.What did not survive contact with the code
Reported rather than quietly dropped:
Pdoes not change published digests, which is stronger than R1 predicted, and the cross-Ptest asserts equality rather than inequality.validate_staged_detailsandreject_staged_base_conflicts— both read the un-surrogated domain and reject a non-zero surrogate as non-canonical. Per-partition identity counts are recorded in the shape intent so S6 can compute the prefix sum without an extra pass.source_ordinalis unnecessary here. Identity UUIDs are globally unique across the staged domain (duplicates are refused), so the UUID alone is already a total order and no tiebreaker is needed. Registration order therefore cannot affect the result, which is what R3 existed to guarantee.ConstructionChunkReceiptkeepssequenceas its primary key and staged chunks are still per-sequence; partitioning happens at shape time, not at intake. Moving it into intake is an S6-shaped restructure.recover_shape_intentunlinks the shape intent and restarts shaping for an incomplete shape, so the recorded splitters are audit and S6 state today. Resume-identical partitioning is instead guaranteed by the sampler being a pure function of the recorded chunk receipts, and is proven byinterrupted_shaping_resumes_to_identical_digests. Preserving the intent across recovery would narrow a crash-safety invariant and is a maintainer call.shape_authority_sha256transitively includesIdentityRecordand is not reproducible across hosts. It is a purely local consistency authority. If anything ever publishes or compares it across hosts, that is a defect.Coordination
main's new commits do not touchgraph_construction. Rebased cleanly onto13632d4b.shape_canonical_inner's per-chunk loop is the one function both changes touch; the S3 authentication passes (authenticate_artifact,validate_parquet_metadata) are left exactly where they were so that removal stays a clean diff.m6_storage_io.rs. The design assigns partition-balance reporting to the benchmark's owner. The evidence fields it needs (shape_partitions,max_partition_identity_rows,partitioned_identity_rows,partition_outputs,peak_partition_records) are now onGraphConstructionEvidence, so that can land without further storage changes.Verification
CI on this head: all required checks pass, including
CI Gate,Windows graphforge-storage Locks(8m44s — every step that was skipped behind theearlier compile error now runs),
macOS graphforge-storage Durability,Native Durability Aggregate,Concurrency MatrixandBazel Bootstrap.Part of #1387
🤖 Generated with Claude Code