Skip to content

observe(storage): extend per-phase I/O attribution beyond the construction path - #1422

Merged
DecisionNerd merged 3 commits into
mainfrom
observe/1389-lifecycle-io-attribution
Sep 17, 2026
Merged

DecisionNerd merged 3 commits into
mainfrom
observe/1389-lifecycle-io-attribution

Conversation

@DecisionNerd

@DecisionNerd DecisionNerd commented Sep 17, 2026 •

Copy link
Copy Markdown
Contributor

What this does

StorageIoPhase attribution existed for the construction path only. This extends it, in the same document shape, to project open, reopen, recount, query, export and clean import, and carries it into ladder rung evidence.

Instrumentation only. No behaviour, scheduling, durability, locking, result or error change.

Mechanism

Modelled on the existing io_stats precedent rather than a parallel accounting scheme:

  • graphforge_storage::lifecycle_io — process-global phase-keyed counters using the same PhaseIoTotals fields and the same {phases, totals} document as ConstructionPhaseAttribution, so existing analysis reads it without new tooling.
  • PhaseScope — a thread-local override for work whose owning phase is only known to the caller (the recovery-on-open pass, which reuses the ordinary generation readers).
  • ReadPathFile — a ChunkReader that delegates entirely to Parquet's own impl ChunkReader for File and only counts, so read placement, buffering and error behaviour are unchanged.

Surfaced as an application_io block on the storage-attribution, result-sink, portable export and portable import CLI receipts; carried through the certify runner behind a fail-closed validator requiring the complete inventory and exact reconciliation; assembled into storage_attribution.lifecycle_application_io in rung evidence.

The one new phase name, and why

read_path_scan. Every existing row names write-side construction, publication, or open-time hydration. None describes decoding already-committed data to answer a read request. Without it, query I/O is either unattributed or mislabelled as hydration — which would destroy exactly the open-versus-execute separation this issue exists to measure.

StorageIoPhase::ALL stays at nine and ConstructionPhaseAttribution still emits exactly those rows. The new row lives only in StorageIoPhase::LIFECYCLE, with its own $defs in the rung and certification schemas, so the existing nine-name phaseMap, the two g500-*.schema.json phase maps and scripts/ci/validate-g500-certification.py are untouched.

Construction attribution is byte-identical — proven by diff

Identical ingest workflow (same scale, seed and operation UUID), run with gf built at the merge base 1c078fa2 and with gf built from this branch, diffing the construction block of the commit receipt:

application_io identical: True
whole construction block identical: True
application_io bytes equal: True (1507 bytes)
baseline sha256: d3c6bf80b0e9f52df9c0be956eee907ef27e40dff51626f6a460439886667902
branch   sha256: d3c6bf80b0e9f52df9c0be956eee907ef27e40dff51626f6a460439886667902

Rung evidence recorded before this change carries the new block nowhere; such a rung omits the field and stays comparable. A run where some lifecycle phases carry it and others do not is refused as a real gap. Both cases are covered by tests.

Measured per-phase attribution

The ordinary gf lifecycle — the same phases and the same commands the ladder profile runs — on a compact v2 project at two scales. Retained graph is 33.6 B/edge (s12) and 34.3 B/edge (s14).

"Opens" counts the project opens performed by the receipts that carry attribution. reopen exceeds 4× because gf storage-attribution also walks the authenticated inventories twice; export and clean_import exceed it because the operation itself re-reads the graph on top of its open.

phase opens s12 read B/edge s14 read B/edge s12 ×retained/open s14 ×retained/open
reopen 2 333.7 342.6 4.96× 4.99×
recount 2 267.1 274.3 3.97× 3.99×
query 2 266.9 274.1 3.97× 3.99×
export 1 233.2 239.7 6.94× 6.98×
clean_import 1 234.6 240.3 6.98× 7.00×
reopen_proof 6 867.7 890.9 4.30× 4.32×

Phase composition of the query phase at s14 (two gf processes, two bounded ORDER BY ... LIMIT traversals):

phase read bytes reads share
hydration_verification 71,814,914 1,292 99.96%
publication_preauthentication 11,220 36 0.02%
read_path_scan 20,404 20 0.03%

For comparison, the existing construction attribution on the same run (unchanged by this PR), s14 ingest: recovery_reauthentication 1,630 B/edge, shape_consume_reauthentication 1,432 B/edge, encode_write_postwrite_authentication 679 B/edge, total 3,812 B/edge read — the same order as the 4,579 B/edge recorded at S22 in #1384, an independent cross-check that the new mechanism agrees with the old one.

The four questions

1. Does opening a project read or verify data proportional to edge count, and in which phase? — Yes, in hydration_verification.

A single project open reads 3.99× the entire retained graph, constant to within 0.5% across a 4× size range. It is hydration_verification at 99.96%. The remaining publication_preauthentication term is flat at 16,252 bytes / 52 reads regardless of graph size — that one is FORMAT, CURRENT, the generation manifest and its participants, and it does not scale.

A control run settles that this is the open and not the query. Four different queries against the same s14 project:

query hydration_verification bytes objects read_path_scan bytes
RETURN 1 (touches no graph data) 35,907,457 138 0
MATCH (n) RETURN count(n) 35,907,457 138 66,419
MATCH ()-[r]->() RETURN count(r) 35,907,457 138 10,202
bounded 1-hop ORDER BY ... LIMIT 1000 35,907,457 138 10,202

A query that reads nothing pays the identical 35,907,457 bytes, to the byte. The hydration figure is the open, in full, and read_path_scan isolates the query's own work correctly.

The code path responsible is verify_graph_object inside ResolvedProjectGeneration::graph_files_inventory() (project_generation.rs:358), whose V2 branch stream-hashes every CAS object, reached from workspace_hydration.rs:394, workspace_hydration.rs:275 and property_overlay/inventory.rs:390, plus a further pass in materialize_graph_objects.

One honest qualification: the counters record 138 object authentications per open against 60 CAS objects — 2.3 per object — while the byte multiplier is 3.99×. The two disagree, so authentication is biased toward the large objects rather than uniform, and the neat "reached three times plus once" reading of the call sites is not established by this instrumentation. Whoever removes the redundancy should confirm the per-object call multiplicity directly; what is established here is the byte volume and that it is authentication.

2. Is any part of the open path authentication or reauthentication, as the construction path is? — Yes; essentially all of it.

Every byte in hydration_verification on the open path is a SHA-256 authentication read. write_bytes is 0 on every open: nothing is copied, CAS objects are hard-linked, so there is no data-movement component to confuse it with. The open path has the same disease as ingest, at 4× rather than 17.3×.

This is not #1269. #1269 (PR #1270, commit 7c0075d9) touched only graph_construction.rs, graph_construction/supersession.rs and construction_lifecycle_tests.rs, and charges recovery_application_read_bytes — it is the recovery_reauthentication row of the ingest attribution, the largest single ingest row at 1,630 B/edge. That work is deliberate and load-bearing. The open-path sweep is a different thing: it dates from #933 ("publish compact authenticated graph roots", project_generation.rs:358), and its multiplier is redundancy — the same objects authenticated four times within one open. Verifying once per open would preserve identical guarantees. The two must not be conflated when #1384 decides what can be removed.

3. Is the cost dominated by I/O or CPU? — Partially answered; see the caveat.

Isolating the terms with gf --info (no project open) as the fixed baseline:

wall user CPU sys CPU attributed reads
gf --info 0.02 s 0.00 0.00 —
gf recovery s12 (1 open) 0.17 s 0.01 0.03 8.75 MB
gf recovery s14 (1 open) 0.27 s 0.03 0.04 35.9 MB

Process startup is only ~25 ms, so the growth is the open. Marginal cost between the two scales is 1.1 ms of CPU per MB authenticated (≈905 MB/s) against ~3.7 ms/MB of wall. So roughly 30% of the marginal cost is on-CPU, and sys exceeds user — the largest single CPU component is read syscalls, not the hash.

Caveat, stated plainly: the remaining off-CPU time could not be cleanly attributed because other agents were building concurrently on this host throughout the measurement window, and the data here is page-cache warm at a scale where the ladder's 0.63 µs/edge regime does not yet apply. I could not settle device-versus-CPU at ladder scale. That needs a rung run carrying this attribution, which this PR enables and I deliberately did not run.

One incidental confirmation for #1384: a storage-attribution invocation authenticates 89.8 MB within 0.07 s of total user CPU, implying a SHA-256 rate above 1.28 GB/s. The host reports sha_ni and OpenSSL measures 1.89 GB/s on it. So the hash primitive is hardware-accelerated, as #1384 assumed, and remains not the lever.

4. Does a bounded query re-pay the open cost per query, or once per session? — Once per session; but the ladder's session is one process per query, so in the ladder it is per query.

In-process, a second identical bounded query on the same handle re-pays zero hydration_verification and zero publication_preauthentication (asserted in crates/graphforge-api/tests/lifecycle_io_attribution.rs). Query execution clones cached Arcs and never re-resolves the generation.

But each ladder query is its own gf process. At s14 the two-query query phase reads 71,846,538 bytes, of which 20,404 bytes (0.03%) is the query's own work and the rest is opening the project twice. A single bounded ORDER BY ... LIMIT 1000 traversal pays 35,907,457 bytes of open authentication to perform 10,202 bytes of its own reads — a factor of 3,520. The query and reopen costs in the ladder are the same cost, counted once per process.

Instrumentation overhead

  • Microbenchmark: 14.3 ns per recorded operation (10M record_read calls in 142.98 ms).
  • The recorder fires once per authenticated object or control file, and once per Parquet read() — roughly 500–2,000 invocations per gf process at these scales, i.e. 7–29 µs against phases of 0.2–3 s.
  • End to end, instrumented against baseline gf, six runs each of storage-attribution on the same project: user 0.06–0.07 s both, sys 0.09–0.11 s both. Not resolvable above noise.

It does not distort what it measures.

Verification

  • cargo fmt --all -- --check clean; cargo clippy --workspace -- -D warnings clean. (--all-targets has pre-existing failures across unrelated crates on main; none in files this PR touches.)
  • graphforge-storage 1267 passed, graphforge-api lib 760 passed, graphforge-cli 110 passed, certify runner 29 passed, plus the new lifecycle_io_attribution integration test.
  • Python harness: 416 tests pass under make smoke-python's discovery (baseline runs 415; this PR adds one).
  • Contract tests updated rather than worked around: the certify receipt sanitizers and their fixtures, the rung and certification-evidence schemas, progressive_run / progressive_storage_qualification inventories, the exhaustive phase_name match in scale_g500_ladder.rs, and the CLI closed-receipt tests.
  • crates/graphforge-cli/tests/portable.rs now asserts the new block separately and compares the remaining package receipt to the facade receipt, so the original "CLI invents or drops no package field" contract is exactly as strict as before.

Known limits

  • The thread-local PhaseScope does not propagate into DataFusion worker threads. Open-path decoding is synchronous on the opening thread, so this is sound today; it is documented in the module.
  • The counters are process-global, so a process that opens two projects concurrently over-reports. The ladder opens one project per process; the facade accessor documents the caveat.

Closes #1389

🤖 Generated with Claude Code


View with [code]smith Autofix with [code]smith
Need help on this PR? Tag @codesmith-bot with what you need. Autofix is disabled.

@coderabbitai

coderabbitai Bot commented Sep 17, 2026 •

Copy link
Copy Markdown

Important

Review skipped

Auto reviews are disabled on this repository. Please check the settings in the CodeRabbit UI or the .coderabbit.yaml file in this repository. To trigger a single review, invoke the @coderabbitai review command.

⚙️ Run configuration

Configuration used: Path: .coderabbit.yaml

Review profile: CHILL

Plan: Advanced

Run ID: df3e8c3c-8de5-4003-a714-c2b6ece136c6

You can disable this status message by setting the reviews.review_status to false in the CodeRabbit configuration file.

Use the checkbox below for a quick retry:

  • 🔍 Trigger review

Warning

Billing warning: we have not been able to collect payment for this subscription for more than 72 hours. Please update the payment method or pay any pending invoices in Billing to avoid service interruption.


Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@github-actions github-actions Bot added core Core source code changes testing Test coverage and testing infrastructure documentation Improvements or additions to documentation ci-cd CI/CD configuration changes tooling Developer tooling and automation labels Sep 17, 2026
@DecisionNerd
DecisionNerd added this pull request to the merge queue Sep 17, 2026
@github-merge-queue
github-merge-queue Bot removed this pull request from the merge queue due to a conflict with the base branch Sep 17, 2026
@DecisionNerd
DecisionNerd force-pushed the observe/1389-lifecycle-io-attribution branch from 10a779f to fff970c Compare September 17, 2026 05:58
@DecisionNerd
DecisionNerd added this pull request to the merge queue Sep 17, 2026
@github-merge-queue
github-merge-queue Bot removed this pull request from the merge queue due to a conflict with the base branch Sep 17, 2026
DecisionNerd and others added 2 commits September 17, 2026 06:43
`StorageIoPhase` attribution existed for the construction path only, which is
why ingest can say 99.1% of its reads are authentication while nothing could be
said about the 30% of lifecycle time spent opening projects. This adds the same
attribution, in the same document shape, to project open, reopen, recount,
query, export and clean import.

Mechanism, modelled on the existing `io_stats` precedent: process-global
phase-keyed counters in `graphforge_storage::lifecycle_io`, a thread-local
`PhaseScope` for work whose owning phase is only known to the caller (the
recovery-on-open pass), and a `ReadPathFile` `ChunkReader` that delegates
entirely to Parquet's own `impl ChunkReader for File` and only counts, so read
placement, buffering and error behaviour are unchanged.

One phase name is added, `read_path_scan`. Every existing row names write-side
construction, publication or open-time hydration; none describes decoding
already-committed data to answer a read request. Without it, query I/O would be
mislabelled as hydration and the open-versus-execute split this issue exists to
measure would be destroyed. `StorageIoPhase::ALL` stays at nine and
`ConstructionPhaseAttribution` still emits exactly those rows, so construction
attribution is byte-identical; the new row lives only in
`StorageIoPhase::LIFECYCLE`, used by `lifecycle_io`.

Surfaced as an `application_io` block on the storage-attribution, result-sink,
portable export and portable import CLI receipts, carried through the certify
runner behind a fail-closed validator that requires the complete inventory and
exact reconciliation, and assembled into rung evidence as
`storage_attribution.lifecycle_application_io`. Evidence recorded before this
change carries the block nowhere; such a rung omits the field and stays
comparable, while a run where some phases carry it and others do not is refused.

Instrumentation only: no behaviour, scheduling, durability, locking, result or
error changes. The emitted document is the closed phase inventory and integer
counters, with no paths, identifiers, query text or property content.

Measured on the ordinary `gf` lifecycle at two scales:

- A project open reads 3.99x the entire retained graph, flat across a 4x size
  range, and 99.97% of it is `hydration_verification`, which is authentication.
- A bounded query's own `read_path_scan` work is 20,404 bytes against
  71,846,538 bytes of re-opening the project.
- A second query in the same session re-pays no open cost at all.
- Instrumentation overhead is 14.3 ns per recorded operation and is not
  resolvable above noise end to end.

Closes #1389

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
… Python

The new integration test crates/graphforge-api/tests/lifecycle_io_attribution.rs
has a Bazel rule but no entry in the fail-closed migration target map, so the
ledger check refused with cargo_target_count mismatch: map=127 cargo=128.

Adds the mapped entry and bumps the count. Also applies ruff format to three
benchmark harness files; those three diffs are formatting only, confirmed
AST-identical to their previous contents.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

ci-cd CI/CD configuration changes core Core source code changes documentation Improvements or additions to documentation testing Test coverage and testing infrastructure tooling Developer tooling and automation

Projects

None yet

Development

Successfully merging this pull request may close these issues.

observe(storage): extend per-phase I/O attribution beyond the construction path

1 participant