Umbrella: GLUTEN-12569 (Spark 4.2 support)
Background
Spark 4.2 restructured Storage-Partitioned Join (SPJ) partitioning. Verified against the Spark jars:
GroupPartitionsExec is new in Spark 4.2.0 (absent in 4.1.1). In 4.2 the scan (BatchScanExec) emits ungrouped partitions — one per split, duplicate keys retained — and grouping is deferred to GroupPartitionsExec. In 4.1 the scan produced the key-grouped layout itself.
Problem
The Spark 4.2 shim (shims/spark42/src/main/scala/org/apache/spark/sql/execution/datasources/v2/BatchScanExecShim.scala, filteredPartitions) still rebuilds the old 4.1 layout via groupBy(InternalRowComparableWrapper(key)), collapsing duplicate-key splits. For keys [A, A, B] it advertises 3 partitions but builds only 2 native RDD partitions; Spark 4.2's GroupPartitionsExec then derives parent indices assuming the ungrouped count → index mismatch / wrong per-key multiplicity.
Scope / why deferred
Only reachable via scans reporting KeyGroupedPartitioning (DSv2 SupportsReportPartitioning). Verified FileScan (Parquet/ORC/CSV) does not implement it, and Delta does not use DSv2 SPJ. The only native consumer is the Iceberg transformer, which is not enabled on Spark 4.2 (dependency unavailable). It is therefore currently dormant / non-triggerable and cannot be behaviorally tested yet.
Required work (when enabling native keyed DSv2 / Iceberg on Spark 4.2)
- Preserve one slot per original sorted split (pad filtered-out splits) and let
GroupPartitionsExec perform grouping — target layout:
original [A1,A2,B1] -> [[A1],[A2],[B1]]
filter removes A2 -> [[A1],[],[B1]]
filter removes A1 and A2 -> [[],[],[B1]]
or explicitly fall native keyed scans back to vanilla Spark until (1) is implemented.
- Add SPJ tests over an actually offloaded scan with partition-count and result assertions (including a nondefault timezone and worker reuse) — belongs with the
gluten-ut/spark42 UT module.
Reference
Deferred from review discussion on #13126.
Umbrella: GLUTEN-12569 (Spark 4.2 support)
Background
Spark 4.2 restructured Storage-Partitioned Join (SPJ) partitioning. Verified against the Spark jars:
GroupPartitionsExecis new in Spark 4.2.0 (absent in 4.1.1). In 4.2 the scan (BatchScanExec) emits ungrouped partitions — one per split, duplicate keys retained — and grouping is deferred toGroupPartitionsExec. In 4.1 the scan produced the key-grouped layout itself.Problem
The Spark 4.2 shim (
shims/spark42/src/main/scala/org/apache/spark/sql/execution/datasources/v2/BatchScanExecShim.scala,filteredPartitions) still rebuilds the old 4.1 layout viagroupBy(InternalRowComparableWrapper(key)), collapsing duplicate-key splits. For keys[A, A, B]it advertises 3 partitions but builds only 2 native RDD partitions; Spark 4.2'sGroupPartitionsExecthen derives parent indices assuming the ungrouped count → index mismatch / wrong per-key multiplicity.Scope / why deferred
Only reachable via scans reporting
KeyGroupedPartitioning(DSv2SupportsReportPartitioning). VerifiedFileScan(Parquet/ORC/CSV) does not implement it, and Delta does not use DSv2 SPJ. The only native consumer is the Iceberg transformer, which is not enabled on Spark 4.2 (dependency unavailable). It is therefore currently dormant / non-triggerable and cannot be behaviorally tested yet.Required work (when enabling native keyed DSv2 / Iceberg on Spark 4.2)
GroupPartitionsExecperform grouping — target layout:gluten-ut/spark42UT module.Reference
Deferred from review discussion on #13126.