Skip to content

[GLUTEN-12569][CORE] Add Spark 4.2 unit test module (gluten-ut/spark42) - #13022

Draft
manoj-ragupathy wants to merge 21 commits into
apache:mainfrom
manoj-ragupathy:feature/42_ut
Draft

manoj-ragupathy wants to merge 21 commits into
apache:mainfrom
manoj-ragupathy:feature/42_ut

Conversation

@manoj-ragupathy

Copy link
Copy Markdown
Contributor

What changes are proposed in this pull request?

Stacked on #13020 → #13021. Targets main, so the diff currently also shows the earlier commits in the stack. Only the top 6 commits belong to this PR. Draft until its parents merge.

⚠️ Please rebase-merge this PR, do not squash

The bulk of this diff is a pure rename of gluten-ut/spark41 → gluten-ut/spark42, kept as its own commit so git log --follow continues to trace each file into its pre-4.2 history. Squashing collapses the rename together with the content edits and destroys rename detection. This follows the commit-history guidance in #11352. I verified git log --follow resolves through the move commit into the original 4.1 history.

Third step of Spark 4.2.x support (#12569): add the gluten-ut/spark42 unit-test module and turn on the corresponding CI jobs.

  • Copy gluten-ut/spark41 → gluten-ut/spark42 as a standalone rename commit (~1850 files, almost entirely mechanical).
  • Adapt the suites that Spark 4.2 broke, and refresh VeloxTestSettings / ClickHouseTestSettings for tests added, removed or renamed between 4.1.1 and 4.2.0.
  • Add the Spark 4.2 UT jobs to velox_backend_x86.yml, matching the existing 4.1 lanes.

Each generated wrapper was checked to extend the correct Gluten trait — #11800 showed that getting this wrong makes suites silently run on vanilla Spark without the plugin, so they pass while testing nothing.

Part of #12569.

How was this patch tested?

This PR is the test module. Full reactor build:

./build/mvn -Pspark-4.2 -Pscala-2.13 -Pjava-17 -Pbackends-velox -Pspark-ut clean install -DskipTests

BUILD SUCCESS — every suite compiles against Spark 4.2.0.

Honest caveat: I could not execute the suites locally (no libgluten.so in my environment), so the actual pass/fail signal has to come from CI. And the 4.2 UT lanes cannot run until the CI image is rebuilt with the Spark 4.2 resources that #13021 adds to the Dockerfile — docker_image.yml only rebuilds on a Sunday cron, so a committer needs to workflow_dispatch it first (also noted on #12569). I expect the first real run to surface failures needing VeloxTestSettings adjustments, and I will iterate on those.

Was this patch authored or co-authored using generative AI tooling?

Generated-by: GitHub Copilot CLI (Claude Opus 5)

MANOJ RAGUPATHY and others added 21 commits September 14, 2026 19:18
Spark 4.2 added parameters to three case classes, which silently breaks
positional extractor patterns:

- CharType gained a collation parameter.
- AppendDataExec / OverwriteByExpressionExec gained tableName and
  transaction.

Match on type (and on the named field where a value is needed) instead,
which is stable across all supported Spark versions.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Copilot-Session: a4b4451c-d099-4965-871d-12781176d082
(cherry picked from commit b2a23c0)
Spark 4.2 hoisted checkAnswer, checkDataset, assertCached and friends out
of QueryTest into a new QueryTestBase trait, which SharedSparkSession now
mixes in. Because GlutenQueryTest declared its own copies of those
members, every suite mixing both inherited conflicting definitions.

Extending QueryTest (an abstract class up to 4.1, a trait in 4.2) makes
Gluten's versions genuine overrides on every supported Spark version, and
removes the duplicated-member conflict.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Copilot-Session: a4b4451c-d099-4965-871d-12781176d082
(cherry picked from commit a0c9f51)
- Spark 4.2 bumps Netty to 4.2.13, which removed
  PlatformDependent.allocateDirectNoCleaner. Use ByteBuffer.allocateDirect,
  which works on every supported version.
- Spark 4.2's spark-catalyst ships a patched copy of
  datasketches ResourceImpl, tripping the ban-duplicate-classes enforcer.
  Both artifacts are 'provided', so nothing extra is packaged; ignore it.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Copilot-Session: a4b4451c-d099-4965-871d-12781176d082
(cherry picked from commit 5d4853d)
…rk 4.2

Spark 4.2 adds `getBinaryView` to `SpecializedGetters`, which makes
`PlaceholderRow` (and any other `InternalRowSparkCompatible` subclass) fail to
compile as non-abstract:

    BatchCarrierRow.scala:110: error: class PlaceholderRow needs to be abstract.
    Missing implementation for member of trait SpecializedGetters:
      def getBinaryView(x$1: Int): org.apache.spark.unsafe.types.BinaryView

Handled with the existing cross-version seam: `SpecializedGettersSparkCompatible`
gains a `getBinaryView` stub and `InternalRowSparkCompatible` overrides it. The
`Nothing` return type conforms to whichever concrete type the running Spark
version declares, so one definition covers the whole 3.x/4.x matrix -- exactly
how getVariant/getGeography/getGeometry are already handled.

Verified with clean builds (stale target/ classes previously masked this):
spark-3.4, spark-3.5, spark-4.0, spark-4.1 and spark-4.2 all BUILD SUCCESS.

(cherry picked from commit 7210d79)
Spark 4.2 moved postDriverMetrics to SupportsCustomDriverMetrics and made
the reported task metrics an explicit argument. Move doPostDriverMetrics
out of the version-agnostic BatchScanExecTransformer into each
BatchScanExecShim so the call site can differ per version.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Copilot-Session: 8ae7d6bd-f561-417a-a635-40fe24dc67a5
Two Spark 4.2 renames that leak into version-agnostic code:

- SampleExec.seed became Option[Long], resolved lazily via resolvedSeed.
- KeyGroupedPartitioning was renamed to KeyedPartitioning.

Both are now hidden behind SparkShims (getSampleSeed,
isKeyGroupedPartitioning) so gluten-substrait and backends-velox stay
version-agnostic.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Copilot-Session: 8ae7d6bd-f561-417a-a635-40fe24dc67a5
Adds the `spark-4.2` Maven profile and a `shims/spark42` module cloned from
`shims/spark41` with identifiers renamed. No behavioural change to existing
Spark versions: `spark-4.2` is opt-in and no other profile is touched.

- root pom.xml: new `spark-4.2` profile (spark.version 4.2.0, arrow 19.0.0 to
  track Spark 4.2's own Arrow, JDK 17+ / Scala 2.13 enforcers); `spark-4.2`
  added to the spark-version requireActiveProfile list.
- shims/pom.xml: `spark-4.2` profile activates module `spark42`.
- shims/spark42: 37 files, renamed spark41 -> spark42 / Spark41 -> Spark42.

Shim selection needs no hardcoded version: SparkShimDescriptor.DESCRIPTOR is
derived from the SPARK_COMPILE_VERSION build property.

This commit does not compile on its own; the Spark 4.2 source-incompatibilities
are fixed in the following commits so that each can be reverted independently.

(cherry picked from commit c4e06e1)
Spark 4.2 reworked `MemoryStream` factory overloads: the
`apply[A](numPartitions: Int)(encoder, sparkSession)` form no longer exists.
The remaining explicit-session overload is
`apply[A: Encoder](sparkSession: SparkSession, numPartitions: Int)`.

Switch the Spark 4.2 shim to that overload; the encoder is supplied by the
existing `A: Encoder` context bound.

(cherry picked from commit 252505b)
Spark 4.2 dropped the trailing `profiler: Option[String]` parameter from both
`PythonUDFRunner.writeUDFs` overloads. Drop the `None` argument in the Spark
4.2 shim. `ArgumentMetadata` itself is unchanged between 4.1 and 4.2.

(cherry picked from commit 5b2558a)
Spark 4.2 reworked the columnar vector API:

- `WritableColumnVector` gains three abstract members —
  `putBytes(int,int,ByteBuffer,int)`, `putShortsFromIntsLittleEndian(...)` and
  `getBytesAsBinaryView(...)`. Added stub overrides to
  `WritableColumnVectorShim` following the file's existing style.
- `GeographyVal` and `GeometryVal` were removed in 4.2 and `ColumnVector`
  instead exposes `getBinaryView` (`ColumnVector.java:297`). `ArrowColumnarArray`
  drops the two geo overrides and delegates `getBinaryView` to `data`.

Verified against the 4.1.1 and 4.2.0 source trees: GeographyVal/GeometryVal
exist only in 4.1.1, BinaryView only in 4.2.0.

These are independent of the SPJ port and were masked until Scala compiled,
since scala-compile-first runs before javac.

(cherry picked from commit 4d14a83)
Spark 4.2 replaced the storage-partitioned-join (SPJ) plumbing on
`BatchScanExec`:

- `StoragePartitionJoinParams` and the `KeyGroupedPartitionedScan` trait were
  removed; the constructor now takes `keyGroupedPartitioning: Option[Seq[Expression]]`.
- `KeyGroupedPartitioning` was replaced by `KeyedPartitioning`, which exposes
  `partitionKeys: Seq[InternalRowComparableWrapper]`, `keyOrdering`,
  `keyRowOrdering` and `toGrouped` instead of `partitionValues`/`expressions`.
- `filteredPartitions` became `Seq[Option[InputPartition]]`, runtime-filter
  pushdown moved to `PushDownUtils.pushRuntimeFilters`, and `postDriverMetrics`
  now requires the task-metrics argument.

Changes are confined to `shims/spark42`; `gluten-core`, `gluten-substrait` and
the backends are untouched. `Spark42Shims` and `BatchScanExecShim` keep the
exact signatures of their spark41 counterparts, so the `keyGroupedPartitioning`
propagation added in apache#12567 (ScanTransformerFactory) needs no change.

KNOWN LIMITATION (Spark 4.2 only, marked DEGRADED in the code):
`getCommonPartitionValues` returns `None`, and `orderPartitions` no longer
applies `joinKeyPositions`, `reducers` or partially-clustered replication.
Spark 4.2 moved these refinements off the scan node into
`EnsureRequirements`/`GroupPartitionsExec`, so there is no scan-level
equivalent to reproduce. The base fully-clustered SPJ path is unaffected.

This is currently unreachable: the only callers of `getCommonPartitionValues`
are `IcebergScanTransformer` and `PaimonScanTransformer`, and neither Iceberg
nor Paimon publishes Spark 4.2 artifacts yet (both 404 on Maven Central, vs 200
for their 4.1 builds); both profiles are also activeByDefault=false. Gluten
never populates `joinKeyPositions`/`reducers` anywhere. To be revisited when
Iceberg/Paimon ship Spark 4.2 support.

(cherry picked from commit 687d477)
Prepares CI for Spark 4.2 without yet adding the UT jobs (those follow with
gluten-ut/spark42):

- install-spark-resources.sh: new `4.2)` case installing Spark 4.2.0. As with
  4.0/4.1, Spark does not publish a `-scala2.13` binary, so the 2.12 tarball is
  installed and `assembly/target/scala-2.12` is renamed to `scala-2.13`. Paths
  are spark42-specific (cf. apache#11973, where the 4.1 case wrongly referenced
  spark40).
- Dockerfile.centos{8,9}-dynamic-build: bake Spark 4.2 resources into the CI
  image, which is where /opt/shims/sparkNN/spark_home comes from.
- velox_backend_x86.yml: add a `shims42` change-detection flag, wired in all
  four required places (outputs block, both scheduled/dispatch flag loops, and
  the path-match line). Missing any one of these makes Spark 4.2 jobs silently
  skip while the PR still reports green.

Verified: spark-4.2.0-bin-hadoop3.tgz and spark-4.2.0.tgz both return HTTP 200
from the mirror the installer actually uses.

(cherry picked from commit 68e19d4)
Mirrors the existing spark-4.1 profile so the micro-benchmark tool can be
built against Spark 4.2.0 / Scala 2.13.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Copilot-Session: a4b4451c-d099-4965-871d-12781176d082
(cherry picked from commit e2540dd)
Implements the Spark 4.2 side of the shim contracts introduced earlier:
doPostDriverMetrics on BatchScanExecShim, plus getSampleSeed and
isKeyGroupedPartitioning on Spark42Shims.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Copilot-Session: 8ae7d6bd-f561-417a-a635-40fe24dc67a5
…ark42

The CI license-header check inspects newly added files. Replace the short
Apache form carried over from shims/spark41 with the canonical ASF header.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Copilot-Session: 8ae7d6bd-f561-417a-a635-40fe24dc67a5
Pure rename so that git preserves file history for the Spark 4.2 test suites.
gluten-ut/spark41 is restored unchanged in the next commit; this commit and the
next are a history-preserving pair.

NOTE FOR MERGE: this PR must be rebase-merged, not squash-merged, or the
rename/restore pair collapses and `git log --follow` no longer works for the
gluten-ut/spark42 files.

(cherry picked from commit 1f06cb8)
Restores gluten-ut/spark41 unchanged after the previous rename commit. The
rename/restore pair makes git attribute the gluten-ut/spark42 tree to the
spark41 history, so `git log --follow` works on the new Spark 4.2 suites while
Spark 4.1 support is left completely untouched.

Must be rebase-merged, not squash-merged.

(cherry picked from commit 52b3021)
Adapts the moved test module to Spark 4.2 and activates it:

- gluten-ut/pom.xml: `spark-4.2` profile activates module `spark42`.
- gluten-ut/spark42/pom.xml: artifactId `gluten-ut-spark42`, name updated.
- GlutenPlanStabilitySuite: golden-file resource dir and the documented
  regeneration command now point at spark42 / `-Pspark-4.2` /
  /opt/shims/spark42/spark_home; the "replace with the previous version"
  hint now refers back to Spark 4.1.

Only 2 of the 1853 files in the tree referenced the Spark version at all, so
the suites themselves are carried over unchanged; Spark 4.2-specific test
adjustments follow separately.

(cherry picked from commit 65926da)
- KeyGroupedPartitioning -> KeyedPartitioning (the suite class itself
  keeps its upstream name); collectShuffles/collectAllShuffles are now
  protected upstream, so widen the overrides.
- BroadcastHashJoinExec gained isSkewJoin and DataSourceV2ScanRelation
  gained pushedFilters; stop matching them positionally.
- HashedRelationSuite and StreamingJoinSuite gained abstract members and
  were split upstream into on-heap/off-heap and VCF/non-VCF variants.
  Mirror that split and register the new suites in VeloxTestSettings.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Copilot-Session: a4b4451c-d099-4965-871d-12781176d082
(cherry picked from commit 014892a)
Adds spark-test-spark42 and spark-test-spark42-slow, cloned from the
spark41 jobs. The -Pdelta lane is omitted because delta-spark does not
yet publish a Spark 4.2 build.

These jobs only pass once the CI container image has been rebuilt with
install-spark-resources.sh 4.2 baked in (dev/docker/Dockerfile.centos*).

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Copilot-Session: a4b4451c-d099-4965-871d-12781176d082
(cherry picked from commit aa236e5)
…t/spark42

The CI license-header check inspects newly added files. Move the ASF header
directly after the XML prologue, where the checker expects it.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Copilot-Session: 8ae7d6bd-f561-417a-a635-40fe24dc67a5
@github-actions

Copy link
Copy Markdown

Run Gluten Clickhouse CI on x86

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant