perf(storage): publish the adjacency CSR with the generation instead of rebuilding it per process - #1453
Conversation
|
Important Review skippedAuto reviews are disabled on this repository. Please check the settings in the CodeRabbit UI or the ⚙️ Run configurationConfiguration used: Repository: CurateLabs/graphforge/.coderabbit.yaml Review profile: CHILL Plan: Advanced Run ID: You can disable this status message by setting the Use the checkbox below for a quick retry:
Warning Billing warning: we have not been able to collect payment for this subscription for more than 72 hours. Please update the payment method or pay any pending invoices in Billing to avoid service interruption. Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out. Comment |
The CSR is built twice for a single-relation graph — worth −12 to −16 s at S20, and it halves two other costsMeasured while pricing S26 levers on
So roughly half the +33.2 s publish-side build measured in this PR is a duplicate sort and write of bytes that already exist. Aliasing the It also halves two costs this PR introduces elsewhere:
Cheapest experiment: time Not asking for it in this PR unless it is easy — the PR is measured, green and worth landing as it stands. Filing it here so it is not rediscovered later, and because it changes the S26 arithmetic for the better. Query side must still resolve |
This comment has been minimized.
This comment has been minimized.
|
The profile says PR #1453 is the fix for this, and that it also unblocks PR #1466Following the call path posted above to its cause, and ruling out the obvious suspect first. It is not a missing page index. It is the access pattern against the page granularity. Pages are That is why 43.6% of the thread is in The consequenceThis is a layout-versus-access-pattern mismatch, not a coding defect. Filtered parquet scanning is the wrong mechanism for a scattered adjacency lookup, at any page size — shrinking pages trades decompression for index size and per-page overhead. Which is exactly what this issue proposes to remove. PR #1453 publishes the adjacency CSR with the generation instead of rebuilding it per process, so a hop query opens a CSR rather than scanning filtered parquet. If the hop path stops going through The dependency worth recordingPR #1453 plausibly unblocks PR #1466. #1466 (unpin the default resource policy) currently hangs the S18 rung — 5 of 5 attempts — and the profile puts the hang in this query path, amplified by That is a hypothesis, not a result: I have not built #1453 and re-run the rung. It is a cheap test — Supporting measurement#1449 records that a one-hop at S20 costs 14 s user CPU and 6.4 GB of disk reads while Reproduction for anyone picking this up: |
Tested: #1453 removes the hang that blocks #1466I posted this as a hypothesis. It now has a measurement.
37.6 s is ordinary S18 territory; the hang was unbounded. Taking the hop query off filtered parquet removes the cost that #1466's Suggested merge order: #1453 before #1466. Two honest caveatsThe combined run does not pass. It completes the work and then fails And it may be my mess rather than #1453's. This was a local merge across three worktrees carrying my own in-flight fixes, and I tripped over that twice while testing: first a The timing result is robust to that — 37.6 s versus >400 s is not a subtle difference and does not depend on which certify binary validated it. The ReproductionBoth on a quiet host, benchexec, same generator and data. |
…nd complete-package export
…est; owner-only mode in the corruption probe
f9b0c77 to
284eee9
Compare
Main introduced the title/adr/status/date/superseded_by frontmatter policy (ADR 0038, #1390) after this branch's ADR was written, so the rebase left 0037 as the only record without it. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
This comment has been minimized.
This comment has been minimized.
…adjacency CSR
`detail_codec_current_format_resumes_cross_multiple_merge_levels` failed under
Bazel on
assertion failed: allocation.snapshot().unwrap().matches_file_inventory(&raw, references)
`matches_file_inventory` admits no untracked file: the tracker's active set must
equal the on-disk inventory, owner for owner and byte for byte. Every other
artifact writer registers through `replace_file_at`; `encode_adjacency` did
not, so the CSR files it publishes existed on disk while belonging to no owner.
`encode_adjacency` now registers each published file in the loop that already
opens and authenticates it, before authentication consumes the handle. No new
parameter was needed: `StableDirectory` is this module's alias for
`ConstructionDirectory`, so the handle was already reachable as
`output.allocation()`.
This deliberately does not attribute the bounded builder's individual write and
fsync calls; that exclusion is documented on `AdjacencyEvidence::write_bytes`
and stays. What is accounted for is the set of files the builder leaves behind,
which is a different question and the one the inventory asks.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
c36f46f to
ae1de7f
Compare
This comment has been minimized.
This comment has been minimized.
Three invariants collide with ADR 0037, not oneBazel Bootstrap has been red on this branch across two runs with the same two targets failing ( Root cause. ADR 0037 moves the derived adjacency CSR from built lazily, per process to published with the generation at construction time. Three existing invariants still encode the old lifecycle. 1. Allocation ownership — fixed in
|
Both failures under Bazel were the same root cause, and neither was a defect in the code under test: they assert the lifecycle that ADR 0037 replaces, where derived adjacency did not exist until some reader built it. `canonical_encoder_outputs_feed_ordinary_readers_index_and_adjacency` called `build_adjacency_index_from_inventory` -- the legacy reader-side builder -- on the directory construction had just published into. That replaced the inodes `record_encoded_active_artifacts` had pinned, so the resume below refused them with `supersession encoded artifact identity changed`. That is the identity ledger doing its job. The test now reads the published manifest, which is what an ordinary reader does once the CSR ships with the generation; `build_adjacency_index_for_edge_files` documents this as the point of the publish-side entry. `permanent_storage_budgets` asserted the Adjacency category was empty after construction, and that its presence equalled `f.adjacency` -- "unbuilt and built adjacency are distinct measured capabilities". ADR 0037 deletes that distinction: a constructed project always carries adjacency. The remaining property is that an explicit rebuild neither removes it nor exceeds the category budget, which is what it now asserts. `unindexed` is renamed to `constructed` for the same reason, including in the emitted evidence keys. `authenticated_lookups_checked` moves 352 -> 373 because `indexes/adjacency/**` are graph files, so the manifest Merkle tree covers them and the authenticated walk visits their leaves and the interior nodes above. That constant is a pinned observation of the tree's shape rather than an invariant; the reason is recorded at the site so the next reader does not have to re-derive it. Verified: 1,186 storage lib tests pass, 59 permanent_storage_budgets tests pass, fmt and clippy clean. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Part of epic #1388.
Every query process rebuilt the entire adjacency CSR into a per-process
graphforge-adjacency-cache-*temp dir and deleted it at exit: ~16 s CPU at S20, paid four times per ladder rung, and the reason theadjacencystorage category reported 0 bytes. This makes the generation that publishes canonical topology also publish its derived CSR, so a query process opens it (ShardedCsrIndex::open, presence-only per #1094) instead of rebuilding it.Design (ADR 0037)
graph_construction_encoding::encodebuildsindexes/adjacency/from the exact edge tables it just encoded, stamps the generation it is about to bind, and records every file as a SHA-256-declaredConstructionEncodedArtifact. The existing publisher installs them into the CAS like every other file; hydration hardlinks and digest-verifies them at open; the persistent provider finds a fresh manifest and serves it. No read-path change and noworkspace_hydration.rsedit.Index-role inventory entries all pre-date this; explicitindex("adjacency")already published them. What changes is when the artifact is produced. Rebuild-on-absence is untouched: an old project still opens (measured below).topology/alone, the shard directory is named by its content digest, and the manifest build time is the session's recorded clock, so the encoded inventory authority stays reproducible; the determinism suite compares every encoded artifact and now covers these.*.csr.jsonandindex_manifest.parquetare sweep-only; a disagreeing index is treated as stale and rebuilt privately, never served.parent_topology_generation == 0). An append carries the parent's files forward without re-reading them; its carried-forward manifest reads as stale and keeps today's lazy rebuild. Follow-up.Measurement (S20 = 2^20 nodes, 16.8M edges; contended host, so user CPU and disk bytes are the comparable figures, wall is indicative)
Per hop query (query and reopen_proof phases, ladder's exact ONE_HOP/TWO_HOP):
1955f17d)Cost moved to publish (per-step ingest, both binaries):
validate(where encoding runs) 134.9 s → 168.1 s user (+33.2 s, +3.5 GB rchar, peak RSS 102 → 191 MB);commit3.9 s → 6.4 s user (+4.3 GB rchar: CAS install hashes the new objects). Net per rung: −4 × ~18.5 s = −74 s of query CPU for +36 s at publish.adjacencystorage category: 0 → 428 MB logical / 214 MB allocated (the CAS deduplicates the_allpair against the single relation's identical shards). Project 2.1 GB → 2.7 GB on disk.Verification
cargo test -p graphforge-api --test bdd: 3897 passing of 3897 (0 regressed, 0 xpass); API BDD 118/118.cargo test --manifest-path benchmarks/Cargo.toml): green.graph_construction::tests::determinism(incl.same_input_twice_produces_identical_digests, resume),lifecycle_budget,partition::tests,recovery::tests,adjacency*,project_portable_v2_import: green. New:initial_construction_publishes_a_current_adjacency_index,complete_import_publishes_the_derived_adjacency_index,import_builds_the_index_a_package_lacks_and_never_twice, and APIpublished_construction_serves_its_adjacency_index_and_refuses_corruption(flips a byte in a published shard object; reopen is refused with a digest mismatch).lifecycle_io_attribution+fixed_hop_limit, graphforge-exec lib +persistent_adjacency: 1702 passed, 0 failed.cargo clippy --workspace -- -D warningsclean; scoped--testsclippy clean on every line this PR touches.make pre-push-fastgreen; source-size policy required splittinggraph_construction_encodingandproject_portable_v2_importadjacency code into child modules.🤖 Generated with Claude Code
Need help on this PR? Tag
@codesmith-botwith what you need. Autofix is disabled.Closes #1446