Repository navigation
Conversation
…rators Adapt Apache SPARK-58968 (apache#58262 and apache#58469) to the branch-4.0 scan architecture with conservative distribution enforcement. Generated-by: OpenAI Codex (version not exposed in this session).
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
JIRA: SPARK-58968. This is the branch-4.0 adaptation of the existing upstream issue.
Draft: standalone and full-build validation remain pending.
What changes were proposed in this pull request?
Adapt SPARK-58968 to branch-4.0. The master fix (#58262) and its 4.2 backport (#58469) use
GroupPartitionsExec, which branch-4.0 does not have. This adaptation inserts a hash shuffle before a single clustered consumer when its storage partitioning needs a key projection. The multi-child join path is unchanged.Full-key and derived-expression clustering remain usable. A
PartitioningCollectionis sufficient when any member meets the actual clustering, and required partition counts still apply.Why are the changes needed?
With subset keys enabled, a scan partitioned by
(k, discard)can report that it satisfies a window'sPARTITION BY k, although equalkvalues are on separate partitions. On rows(1,10,10)and(1,20,20),ROW_NUMBER() OVER (PARTITION BY k ORDER BY v)can therefore return1,1instead of1,2. The grouping must be established before the window executes.Does this PR introduce any user-facing change?
Yes. Single-child operators requiring clustered input produce correct results when their keys are a strict subset of the storage partition keys.
This is a conservative adaptation, not a literal cherry-pick. Unlike the newer grouping implementation, it can add a hash shuffle even when projecting the reported key values would leave every key distinct. It avoids rewriting a scan beneath an operator that depends on its existing distribution or ordering.
How was this patch tested?
Added window regressions for standalone and joined queries, both values of
requireAllClusterKeysForDistribution, and full-key controls. The tests check row numbers and require the subset shuffle to be belowWindowExec. Added planner coverage for transformed keys, mixed partitioning collections, and required partition counts.git diff --checkpassed.Baseline controls returned duplicate row numbers in the standalone case and lacked the required shuffle below the joined window. Both candidate cases passed, including their full-key controls.
Fresh selected-source local validation of a combined branch-4.0 tree containing the separate window, key-routing, scan-ordering and NULL-fixture proposals passed all 11 focused tests. The broader run completed all 244 exact test identities: 237 passed and seven function-based ordering cases in WriteDistributionAndOrderingSuite failed. All seven also failed on the Apache production baseline with matching exception types, messages and first eight stack frames; they remain a limitation of this local setup, not a passing full-suite result. No tests were skipped and no suite aborted; fresh-class origin checks passed.
The separate baseline control reproduced six targeted production regressions, while five expected controls passed. Removing only the two NULL-fixture guards separately reproduced the insertion exception and stale-value result. Candidate tests cover their complete loops; baseline failures stop at the first failing iteration and do not establish later iterations. All nine changed files in the combined 4.0 tree passed Scalastyle with zero errors or warnings.
JDK 17 / Scala 2.13.16 freshly compiled 40 selected Scala production sources, five Java sources and nine test/fixture sources; remaining dependencies were cached and fingerprinted. This is not a complete build, standalone validation of this PR, a whole-source-equivalent Apache runtime, or CI success. Earlier failed setup and audit attempts are retained separately and are not counted as passes.
Was this patch authored or co-authored using generative AI tooling?
Generated-by: OpenAI Codex (version not exposed in this session).