Skip to content

fix(storage): compress permanent construction Parquet with Zstd #1202

Description

@DecisionNerd

Maintainer scope: before v1.0.0

Backward compatibility is out of scope for this issue and its repair sub-issues before the v1.0.0 release. Do not add or retain legacy readers, old API aliases, compatibility shims, mixed-version support, migration/backfill machinery, or old-version regression tests solely to preserve behavior or data from earlier GraphForge versions. Earlier backward-compatibility requirements in this issue are superseded by this maintainer instruction.

This does not authorize unrelated breaking changes or weaken the issue's current-version acceptance criteria. Exact current-format semantics, supported current-version API/binding and export/import interoperability, authentication, corruption/unsupported-format refusal, crash recovery, retry/idempotency, active snapshots, cancellation and resource budgets remain required where applicable. Never silently reinterpret unsupported old data. Document intentional format/API breaks and the supported current format; a migration implementation is not required. Current-version correctness tests and explicitly behavior-preserving refactors remain in scope. Compatibility guarantees for v1.0.0 and later are a separate release-policy decision.

Problem and evidence

#1196's exact Arrow-batch experiment finds construction's permanent Parquet writer uses uncompressed output. On 65,537-edge fixtures, Zstd level 1 changes 5,318,519→4,015,814 bytes for random UUID property-free topology, 10,231,116→7,845,372 for eight-route properties, and 8,357,272→6,398,839 for heterogeneous properties. Sequential UUIDs compress more strongly but are not the acceptance proxy for random identity data.

Scope and acceptance

  • Compress new permanent construction Parquet payloads with Zstd level 1, preserving all UUIDs, full-width IDs, runtime catalog/ontology distinctions, schema/null values, durability, SHA authentication and existing bounded row-group/cache behavior.
  • Current-format reopen, exact query, export/full-verify/clean-import remain identical. Backward compatibility and migration are not required (maintainer instruction: pre-v1).
  • Deterministic fixtures from test(storage): establish permanent topology and identity compaction budgets #1196 prove new Parquet bytes no greater than 80% of their uncompressed baseline for each random-ID fixture, with physical allocation reported separately. Sequential savings are additional evidence only.
  • Existing construction cancellation/recovery, active snapshot, bounded I/O/RSS and ladder resource budgets pass; measure and document encode/decode tradeoffs without treating noisy wall-clock timing as a hard correctness assertion.
  • Focused PR, exact-head required CI/CI Gate and appropriate local checks pass; document the codec policy and source-bound before/after evidence.

Given a current-format project, when a compressed child is published and reopened, then active snapshots retain exact public semantics and portable verification. Given interrupted encoding, recovery authenticates only complete payloads under existing resource bounds.

Native child of #1194, blocked by assessment #1196; blocks #1194. This isolates one validated material repair. Excludes membership, CSR and manifest format changes, unrelated M3 work and S24/S26 certification. Debt: representation cost; Rust remains the engine.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions