Repository navigation
Conversation
Defer native scan DPP partition enumeration until execution so AQE can coalesce a sibling shuffle without executing an adaptive placeholder. Pass codegen scalar subqueries as runtime inputs and broadcast length-one Arrow arguments, including the null fast path. Normalize Etc/UTC in native timestamp truncation to match scan timestamp types. Scale sort's eager spill reserve with small off-heap task budgets. Spill whole-partition aggregate window rows while retaining the existing native accumulators and their Spark semantics. Add regressions for planning, multi-row scalar inputs, timezone aliases, reservation sizing, window spilling, ordering and cancellation cleanup.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Which issue does this PR close?
No issue is closed automatically. This addresses the failures behind the feature overrides in datasci's SparkSessionProducer.
Rationale for this change
Warehouse jobs disable native sort, windows, codegen expressions, timestamp truncation and AQE coalescing to avoid failures. Those overrides cause Spark operator fallback and prevent jobs from benefiting fully from Comet. This patch fixes the reproduced failures while retaining Comet execution and the existing memory limits.
For example, a scalar subquery used as a regexp pattern must produce both
123and456for a two-row batch; it must neither serialize an unresolved subquery nor read past its single-value Arrow argument. A whole-partition aggregate over a large, wide input must spill its rows instead of retaining the entire partition in memory.What changes are included in this PR?
Etc/UTCalias in native timestamp truncation so its output type matches native scan timestamps.sum,avg,count,minandmaxwindows using the existing native accumulators. Other window shapes keep their existing implementation.concat_ws(array<string>, ...)already works in the current upstream base and needs no additional source change.The branch is based on
aaa33eb203422fa2c1015740e5d889825dc48985. It does not change datasci configuration or deploy a new JAR. Window spilling is scoped to the aggregate frames above; this is not general spill support for every unbounded window function. Production cost savings have not been measured.How are these changes tested?
All execution is local in Docker using Colima, with Spark 3.5.3, Java 17 and Linux ARM64.
-D warnings, Rustfmt andgit diff --checkpassed.Spark SQL gate passed:
dev/local-ci.sh spark 3.5 sql_core-1on Spark 3.5.9 / Java 17 completed against8a75822fewith 9,205 passed, 0 failed, 590 suites completed, 0 aborted; 11 canceled and 246 ignored. Total runtime was 24m13s, including preparation. Previously failing zero-file scan cases pass in the broad suite too. The local run used a 6-CPU / 7-GiB dev container and capped only the SBT coordinator heap at 1 GiB; SQL test selection and test-JVM settings were unchanged. No additional cgroup OOM occurred.Canceled-test follow-up: six of the 11 cancellations came from missing pandas in the local test environment. After installing numpy 1.26.4, pandas 2.2.3 and pyarrow 12.0.1 within Spark 3.5.9's dependency constraints, those exact six tests were rerun with Comet: 6 passed, 0 failed, 0 canceled across QueryCompilationErrorsSuite, PythonUDTFSuite and PythonUDFSuite. This is a separate targeted run, in addition to the 9,205-test broad gate.
The other five cancellations are the ORC/Parquet/CSV/JSON/text variants of
ignoreMissingFiles. They reach an unconditionalassume(false)in the pre-existingdev/diffs/3.5.9.diffpatch, before an assertion requiringSparkFileNotFoundException. The upstream comment concerns native Parquet, but the shared test loop cancels all five formats. This PR does not modify that patch, and those five exception checks have not been established as passing.The PR remains draft as requested. Validation covers the reproduced failures and this Spark SQL matrix row, not every supported Spark version or production cost savings.