You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
The conservative Parquet runtime-filter schema guard in #6067 preserves conversion errors, but it also disables reader pruning for adaptations that can be proved infallible and adds work for fully eligible wide scans.
The focused native release benchmark at c48d39f07, documented in PR #6067, found:
Case, runtime filtering enabled
Before guard
With guard
Reader work
Allowed INT32 to BIGINT promotion
1.11 ms
15.36 ms
1 to 128 row groups; about 128 times the data bytes
Struct projection, two of eight fields
1.33 ms
15.92 ms
1 to 128 row groups; about 128 times the data bytes
Fully eligible wide selective scan
29.64 ms
35.72 ms (+20.5%)
Same bytes and row groups
The promotion and nested fixtures use 16 files and 128 row groups. The wide fixture uses 128 small files and 64 columns. Measurements use warm local caches, the standard optimized release profile, and 14 measured samples per case/mode. Timing includes native planning and execution. Nested column pruning remains effective; the struct case loses row-group pruning.
Describe the potential solution
Establish which schema adaptations are infallible and allow reader filtering through those adaptations. Start with permitted INT32-to-BIGINT promotion and projection of a subset of a struct's fields. Use an explicit conversion-safety contract so the decision stays aligned with the existing Spark adapter.
Avoid one inner.rewrite call per required column per file when the logical and physical schemas already match and adapter semantics establish that the fast path is safe. Keep unresolved-column handling, partition exclusions, case and field-ID resolution, and missing/default-column behavior correct.
Keep unresolved names ineligible and preserve the partition and column-resolution regressions.
Preserve static pruning, missing/default values, and supported name/field-ID mappings.
Compare the changes using the same direct, promotion, nested, wide selective, and wide non-pruning fixtures with runtime filtering both enabled and disabled. Report query time, requested data bytes, and row groups read.
Additional context
Requested in this review. This follow-up covers the schema guard's pruning eligibility and per-file checking cost. #5775 tracks TopK fusion and reader pushdown.
What is the problem the feature request solves?
The conservative Parquet runtime-filter schema guard in #6067 preserves conversion errors, but it also disables reader pruning for adaptations that can be proved infallible and adds work for fully eligible wide scans.
The focused native release benchmark at
c48d39f07, documented in PR #6067, found:The promotion and nested fixtures use 16 files and 128 row groups. The wide fixture uses 128 small files and 64 columns. Measurements use warm local caches, the standard optimized release profile, and 14 measured samples per case/mode. Timing includes native planning and execution. Nested column pruning remains effective; the struct case loses row-group pruning.
Describe the potential solution
inner.rewritecall per required column per file when the logical and physical schemas already match and adapter semantics establish that the fast path is safe. Keep unresolved-column handling, partition exclusions, case and field-ID resolution, and missing/default-column behavior correct.Acceptance criteria:
Additional context
Requested in this review. This follow-up covers the schema guard's pruning eligibility and per-file checking cost. #5775 tracks TopK fusion and reader pushdown.