Skip to content

fix(storage): unify permanent Parquet encoding across publishing paths - #1227

Merged
DecisionNerd merged 4 commits into
mainfrom
fix/1213-permanent-parquet-policy
Sep 10, 2026
Merged

DecisionNerd merged 4 commits into
mainfrom
fix/1213-permanent-parquet-policy

Conversation

@DecisionNerd

@DecisionNerd DecisionNerd commented Sep 10, 2026 •

Copy link
Copy Markdown
Contributor

Description

Construction publishes Zstd Parquet, but mutation and replay/compaction could publish uncompressed replacements. Apply one explicit Zstd level 1 encoding policy to all 16 verified permanent-writer sites, including projections, semantic composition, runtime catalogs, vectors, knowledge/provenance and checkpoint restoration. Private construction streams and external result exports remain outside permanent graph storage.

Replay retains disabled dictionaries and its row-group cap; staging retains 65,536-row groups. Account for native codecs, decoder pages/returned batches, schemas and retained writer metadata before admitting replay. When simultaneous node decoding/encoding exceeds the existing 2 MiB regression budget, an independently admitted, self-deleting Arrow stream separates the phases, with a checked 64 MiB per-invocation disk ceiling. Writer lifecycle, authentication and publication implementations remain separate.

Compaction Parquet falls from 537,467 to 303,718 bytes (allocated: 569,344 → 335,872). The full lifecycle observed 179.6 → 163.1 MB syscall reads and 42.7 → 40.2 MB writes; elapsed time rose 12.48 → 12.99 s and peak RSS 156,284 → 159,492 KiB. Measurements, limitations and deterministic budgets are recorded in docs/book/architecture/permanent-storage-assessment.md.

Related Issues

Fixes #1213. Part of #1194; does not close the epic.

Evidence and tests

The assessment and source-bound JSON record current-versus-candidate codec experiments, permanent allocation, CPU/RSS, syscall I/O and sampled workspace peaks. Deterministic tests cover output bytes/allocation, actual column codecs, replay dictionary/row-group settings, exact resource-limit boundaries and private-stream cleanup. Codec compression is not presented as a whole-process memory proof.

  • Public publishing budget suite: 20 passed, including construction/mutation/compaction, active snapshots, reopen/query/export/full verify/clean import and heterogeneous/qualified routes.
  • API unit suite: 734 passed, including recovery/cancellation and published projection/vector/knowledge/provenance/restoration metadata.
  • Public semantic-composition certification: 3 passed, including retained-data migration and cancellation.
  • Storage compaction integration: 10 passed; journal integration: 16 passed; replay-focused unit tests: 20 passed.
  • Targeted Bazel semantic-certification/journal/compaction targets: all passed.
  • cargo clippy --workspace -- -D warnings, make pre-push-fast, make gate-registry-check, Cargo/Bazel drift and final formatting passed.
  • Local full storage aggregate: 1,081 passed, two existing ignores; one unchanged test hardcodes /tmp, where this host's tmpfs fails filesystem admission. Required native Bazel CI remains the merge authority.

Independent read-only review verified the production inventory, resource composition and lifecycle changes; no actionable finding remained. This is one coupled encoding/resource repair; most of the diff is admission accounting, regression coverage and measured evidence. Backward compatibility and format migration machinery are out of scope before v1.0.0.


View with [code]smith Autofix with [code]smith
Need help on this PR? Tag @codesmith-bot with what you need. Autofix is disabled.

@coderabbitai

coderabbitai Bot commented Sep 10, 2026

Copy link
Copy Markdown

Important

Review skipped

Auto reviews are disabled on this repository. Please check the settings in the CodeRabbit UI or the .coderabbit.yaml file in this repository. To trigger a single review, invoke the @coderabbitai review command.

⚙️ Run configuration

Configuration used: Path: .coderabbit.yaml

Review profile: CHILL

Plan: Advanced

Run ID: 8f83c878-43e2-49e6-9b43-d819c1b3f2c2

You can disable this status message by setting the reviews.review_status to false in the CodeRabbit configuration file.

Use the checkbox below for a quick retry:

  • 🔍 Trigger review

Warning

Billing warning: we have not been able to collect payment for this subscription for more than 72 hours. Please update the payment method or pay any pending invoices in Billing to avoid service interruption.


Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@github-actions github-actions Bot added core Core source code changes documentation Improvements or additions to documentation labels Sep 10, 2026
@DecisionNerd
DecisionNerd merged commit ec77585 into main Sep 10, 2026
23 checks passed
@DecisionNerd
DecisionNerd deleted the fix/1213-permanent-parquet-policy branch September 10, 2026 08:18
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

core Core source code changes documentation Improvements or additions to documentation

Projects

None yet

Development

Successfully merging this pull request may close these issues.

fix(storage): unify permanent encoding policy across all Parquet publishing paths

1 participant