fix(storage): unify permanent Parquet encoding across publishing paths - #1227
Conversation
|
Important Review skippedAuto reviews are disabled on this repository. Please check the settings in the CodeRabbit UI or the ⚙️ Run configurationConfiguration used: Path: .coderabbit.yaml Review profile: CHILL Plan: Advanced Run ID: You can disable this status message by setting the Use the checkbox below for a quick retry:
Warning Billing warning: we have not been able to collect payment for this subscription for more than 72 hours. Please update the payment method or pay any pending invoices in Billing to avoid service interruption. Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out. Comment |
Description
Construction publishes Zstd Parquet, but mutation and replay/compaction could publish uncompressed replacements. Apply one explicit Zstd level 1 encoding policy to all 16 verified permanent-writer sites, including projections, semantic composition, runtime catalogs, vectors, knowledge/provenance and checkpoint restoration. Private construction streams and external result exports remain outside permanent graph storage.
Replay retains disabled dictionaries and its row-group cap; staging retains 65,536-row groups. Account for native codecs, decoder pages/returned batches, schemas and retained writer metadata before admitting replay. When simultaneous node decoding/encoding exceeds the existing 2 MiB regression budget, an independently admitted, self-deleting Arrow stream separates the phases, with a checked 64 MiB per-invocation disk ceiling. Writer lifecycle, authentication and publication implementations remain separate.
Compaction Parquet falls from 537,467 to 303,718 bytes (allocated: 569,344 → 335,872). The full lifecycle observed 179.6 → 163.1 MB syscall reads and 42.7 → 40.2 MB writes; elapsed time rose 12.48 → 12.99 s and peak RSS 156,284 → 159,492 KiB. Measurements, limitations and deterministic budgets are recorded in
docs/book/architecture/permanent-storage-assessment.md.Related Issues
Fixes #1213. Part of #1194; does not close the epic.
Evidence and tests
The assessment and source-bound JSON record current-versus-candidate codec experiments, permanent allocation, CPU/RSS, syscall I/O and sampled workspace peaks. Deterministic tests cover output bytes/allocation, actual column codecs, replay dictionary/row-group settings, exact resource-limit boundaries and private-stream cleanup. Codec compression is not presented as a whole-process memory proof.
cargo clippy --workspace -- -D warnings,make pre-push-fast,make gate-registry-check, Cargo/Bazel drift and final formatting passed./tmp, where this host's tmpfs fails filesystem admission. Required native Bazel CI remains the merge authority.Independent read-only review verified the production inventory, resource composition and lifecycle changes; no actionable finding remained. This is one coupled encoding/resource repair; most of the diff is admission accounting, regression coverage and measured evidence. Backward compatibility and format migration machinery are out of scope before v1.0.0.
Need help on this PR? Tag
@codesmith-botwith what you need. Autofix is disabled.