Skip to content

[CORE] Deprecate and remove Spark 3.2 support - #11351

Merged
zhouyuan merged 7 commits into
apache:mainfrom
QCLyu:qingchuanlyu
Jan 20, 2026
Merged

zhouyuan merged 7 commits into
apache:mainfrom
QCLyu:qingchuanlyu

Conversation

@QCLyu

@QCLyu QCLyu commented Jan 4, 2026 •

Copy link
Copy Markdown
Contributor

What changes are proposed in this pull request?

This PR comprehensively removes Spark 3.2 support from the Gluten Velox backend. It cleans up the source code, build profiles, CI/CD pipelines, and documentation.

Key changes include:

  • Source Code: Removed shims/spark32 and gluten-ut/spark32 directories.

  • Build System: Deleted the spark-3.2 profile from the root and all sub-module pom.xml files.

  • CI/CD: Removed legacy Spark 3.2 jobs (spark-test-spark32, spark-test-spark32-slow) from GitHub Workflows.

  • Test Migration: Refactored VeloxHashJoinSuite and other backend tests to remove Spark 3.2-specific conditional logic, ensuring these tests now run on Spark 3.3+.

  • Documentation: Updated the build guide and ClickHouse deployment docs to remove references to Spark 3.2.

How was this patch tested?

  • Manual Build: Verified successful compilation on aarch64 (ARM64) using -Pspark-3.5 -Pbackends-velox.

  • Unit Tests: Verified that migrated tests in VeloxHashJoinSuite pass successfully under Spark 3.5.

  • CI: Infrastructure changes have been validated to ensure remaining Spark versions (3.3, 3.4, 3.5) trigger correctly.

Closes #8960

@github-actions

github-actions Bot commented Jan 4, 2026

Copy link
Copy Markdown

Run Gluten Clickhouse CI on x86

1 similar comment
@github-actions

github-actions Bot commented Jan 4, 2026

Copy link
Copy Markdown

Run Gluten Clickhouse CI on x86

@QCLyu

QCLyu commented Jan 4, 2026 •

Copy link
Copy Markdown
Contributor Author

I have verified the changes for Spark 3.5 locally, while GitHub Actions was showing failures:

Jenkins (ClickHouse CI): SUCCESS.

The Reactor Summary confirms all modules (including gluten-core, shims, and backends-clickhouse) built successfully and all 36 tests passed.

Click to view Jenkins Reactor Summary (Build Success)

12:33:20 Run completed in 2 minutes, 25 seconds.
12:33:20 Total number of tests run: 36
12:33:20 Suites: completed 2, aborted 0
12:33:20 Tests: succeeded 36, failed 0, canceled 0, ignored 12, pending 0
12:33:20 All tests passed.
12:33:21 [INFO] ------------------------------------------------------------------------
12:33:21 [INFO] Reactor Summary for Gluten Parent Pom 1.6.0-SNAPSHOT:
12:33:21 [INFO]
12:33:21 [INFO] Gluten Parent Pom .................................. SUCCESS [ 17.421 s]
12:33:21 [INFO] Gluten Ras ......................................... SUCCESS [ 23.776 s]
12:33:21 [INFO] Gluten Ras Common .................................. SUCCESS [ 51.164 s]
12:33:21 [INFO] Gluten Core ........................................ SUCCESS [ 37.484 s]
12:33:21 [INFO] Gluten Shims ....................................... SUCCESS [ 0.260 s]
12:33:21 [INFO] Gluten Shims Common ................................ SUCCESS [ 6.371 s]
12:33:21 [INFO] Gluten Shims for Spark 3.3 ......................... SUCCESS [ 13.796 s]
12:33:21 [INFO] Gluten UI .......................................... SUCCESS [ 4.449 s]
12:33:21 [INFO] Gluten Substrait ................................... SUCCESS [ 59.459 s]
12:33:21 [INFO] Gluten Celeborn .................................... SUCCESS [ 4.854 s]
12:33:21 [INFO] Gluten Iceberg ..................................... SUCCESS [ 11.367 s]
12:33:21 [INFO] Gluten DeltaLake ................................... SUCCESS [ 9.760 s]
12:33:21 [INFO] Gluten Package ..................................... SUCCESS [ 6.663 s]
12:33:21 [INFO] Gluten Ras Planner ................................. SUCCESS [ 1.137 s]
12:33:21 [INFO] Gluten Kafka ....................................... SUCCESS [ 9.403 s]
12:33:21 [INFO] Gluten Backends ClickHouse ......................... SUCCESS [59:48 min]
12:33:21 [INFO] Gluten Unit Test Parent ............................ SUCCESS [ 1.323 s]
12:33:21 [INFO] Gluten Unit Test Common ............................ SUCCESS [ 5.468 s]
12:33:21 [INFO] Gluten Unit Test ................................... SUCCESS [ 16.964 s]
12:33:21 [INFO] Gluten Unit Test Spark33 ........................... SUCCESS [02:45 min]
12:33:21 [INFO] ------------------------------------------------------------------------
12:33:21 [INFO] BUILD SUCCESS
12:33:21 [INFO] ------------------------------------------------------------------------
12:33:21 [INFO] Total time: 01:07 h
12:33:21 [INFO] Finished at: 2026-01-04T04:33:21Z
12:33:21 [INFO] ------------------------------------------------------------------------

GitHub Actions:

These jobs are failing with 403 Forbidden errors during dependency resolution (Log4j, ASM, etc.). In the following commit, these issues are found and fixed:

  • .github/workflows/velox_nightly.yml (2 occurrences)
    Removed mvn clean install -Pspark-3.2 from both the x86 and arm64 build jobs
    Lines 103 and 226
  • .github/workflows/build_bundle_package.yml
    Updated description from 'Spark version: spark-3.2, spark-3.3, spark-3.4 or spark-3.5' to 'Spark version: spark-3.3, spark-3.4, spark-3.5 or spark-4.0'
  • .github/workflows/util/install-spark-resources.sh
    Removed the Spark 3.2 case (lines 92-96) from the script

These changes aim to resolve the 403 errors. The workflows will no longer attempt to build with the removed Spark 3.2 profile.

@github-actions

github-actions Bot commented Jan 4, 2026

Copy link
Copy Markdown

Run Gluten Clickhouse CI on x86

@github-actions

github-actions Bot commented Jan 5, 2026

Copy link
Copy Markdown

Run Gluten Clickhouse CI on x86

2 similar comments
@github-actions

github-actions Bot commented Jan 6, 2026

Copy link
Copy Markdown

Run Gluten Clickhouse CI on x86

@github-actions

github-actions Bot commented Jan 6, 2026

Copy link
Copy Markdown

Run Gluten Clickhouse CI on x86

@QCLyu

QCLyu commented Jan 6, 2026

Copy link
Copy Markdown
Contributor Author

Hi @zhouyuan , the code changes here look good. The only ClickHouse CI failure is due to the Jenkins job still running -Pspark-3.2 on Java 8, but this repo no longer has a spark-3.2 profile and pulls Iceberg artifacts built for Java 11 (class version 55). That mismatch causes the compile error in gluten-iceberg. This is a CI config issue, not a code regression.

Please proceed with merge if everything else looks good to you, and we can update/disable the Spark 3.2 Java 8 leg in the Jenkins pipeline separately.

This is the only failed step in ClickHouse CI: https://opencicd.kyligence.com/job/gluten/job/gluten-ci/18337/flowGraphTable/
This is the log: https://opencicd.kyligence.com/job/gluten/job/gluten-ci/18337/execution/node/235/log/

18:00:32  [ERROR] COMPILATION ERROR : 
18:00:32  [INFO] -------------------------------------------------------------
18:00:32  [ERROR] /home/jenkins/agent/workspace/gluten/gluten-ci/ut-stage-1/gluten-iceberg/src/main/java/org/apache/gluten/connector/write/MetricsWrapper.java:[21,25] error: cannot access Metrics
18:00:32    bad class file: /root/.m2/repository/org/apache/iceberg/iceberg-spark-runtime-3.5_2.12/1.10.0/iceberg-spark-runtime-3.5_2.12-1.10.0.jar(org/apache/iceberg/Metrics.class)
18:00:32      class file has wrong version 55.0, should be 52.0
18:00:32      Please remove or make sure it appears in the correct subdirectory of the classpath.
18:00:32  [INFO] 1 error
18:00:32  [INFO] -------------------------------------------------------------
18:00:32  [INFO] ------------------------------------------------------------------------
18:00:32  [INFO] Reactor Summary for Gluten Parent Pom 1.6.0-SNAPSHOT:
18:00:32  [INFO] 
18:00:32  [INFO] Gluten Parent Pom .................................. SUCCESS [ 21.426 s]
18:00:32  [INFO] Gluten Ras ......................................... SUCCESS [ 22.345 s]
18:00:32  [INFO] Gluten Ras Common .................................. SUCCESS [ 51.109 s]
18:00:32  [INFO] Gluten Core ........................................ SUCCESS [ 39.051 s]
18:00:32  [INFO] Gluten Shims ....................................... SUCCESS [  0.336 s]
18:00:32  [INFO] Gluten Shims Common ................................ SUCCESS [  6.173 s]
18:00:32  [INFO] Gluten Shims for Spark 3.5 ......................... SUCCESS [ 12.425 s]
18:00:32  [INFO] Gluten UI .......................................... SUCCESS [  5.105 s]
18:00:32  [INFO] Gluten Substrait ................................... SUCCESS [01:01 min]
18:00:32  [INFO] Gluten Celeborn .................................... SUCCESS [  4.345 s]
18:00:32  [INFO] Gluten Iceberg ..................................... FAILURE [  5.848 s]
18:00:32  [INFO] Gluten DeltaLake ................................... SKIPPED
18:00:32  [INFO] Gluten Package ..................................... SKIPPED
18:00:32  [INFO] Gluten Ras Planner ................................. SKIPPED
18:00:32  [INFO] Gluten Kafka ....................................... SKIPPED
18:00:32  [INFO] Gluten Backends ClickHouse ......................... SKIPPED
18:00:32  [INFO] Gluten Unit Test Parent ............................ SKIPPED
18:00:32  [INFO] Gluten Unit Test Common ............................ SKIPPED
18:00:32  [INFO] Gluten Unit Test ................................... SKIPPED
18:00:32  [INFO] ------------------------------------------------------------------------
18:00:32  [INFO] BUILD FAILURE
18:00:32  [INFO] ------------------------------------------------------------------------
18:00:32  [INFO] Total time:  03:49 min
18:00:32  [INFO] Finished at: 2026-01-06T02:00:32Z
18:00:32  [INFO] ------------------------------------------------------------------------
18:00:32  [WARNING] The requested profile "spark-3.2" could not be activated because it does not exist.
18:00:32  [ERROR] Failed to execute goal org.apache.maven.plugins:maven-compiler-plugin:3.14.1:compile (default-compile) on project gluten-iceberg: Compilation failure
18:00:32  [ERROR] /home/jenkins/agent/workspace/gluten/gluten-ci/ut-stage-1/gluten-iceberg/src/main/java/org/apache/gluten/connector/write/MetricsWrapper.java:[21,25] error: cannot access Metrics
18:00:32  [ERROR]   bad class file: /root/.m2/repository/org/apache/iceberg/iceberg-spark-runtime-3.5_2.12/1.10.0/iceberg-spark-runtime-3.5_2.12-1.10.0.jar(org/apache/iceberg/Metrics.class)
18:00:32  [ERROR]     class file has wrong version 55.0, should be 52.0
18:00:32  [ERROR]     Please remove or make sure it appears in the correct subdirectory of the classpath.
18:00:32  [ERROR] 
18:00:32  [ERROR] -> [Help 1]

@zhouyuan

zhouyuan commented Jan 6, 2026

Copy link
Copy Markdown
Member

@zzcclp could you please help to take a look?

@QCLyu

QCLyu commented Jan 7, 2026

Copy link
Copy Markdown
Contributor Author

Hi @zhouyuan @zzcclp Just following up on my last comment here. I’d love to get your thoughts on CI config issue—do you feel this makes sense, or maybe not? Want to make sure we’re aligned before I move forward.

@zzcclp

zzcclp commented Jan 7, 2026

Copy link
Copy Markdown
Contributor

Sorry for the late reply, I modified the CI script , please hava a try again.

@github-actions

github-actions Bot commented Jan 8, 2026

Copy link
Copy Markdown

Run Gluten Clickhouse CI on x86

@github-actions

github-actions Bot commented Jan 8, 2026

Copy link
Copy Markdown

Run Gluten Clickhouse CI on x86

@github-actions

github-actions Bot commented Jan 8, 2026

Copy link
Copy Markdown

Run Gluten Clickhouse CI on x86

@philo-he philo-he left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thank you so much for your work. Some comments. Please check if they make sense.

Suggest to use git pull --rebase <remote name> main to rebase code. Otherwise, some commits already merged to main branch are included in this PR, which are mixed with your changes and not friendly to reviewers.

Please also clean the use of Spark 3.2 in build scripts under dev.

Maybe, some Spark shim APIs were introduced to adapt to the differences between Spark 3.2 and later Spark versions. If so, we should also remove them (can be done in separate PRs).

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I assume we should only remove Spark 3.2 UT test jobs from CI. Other jobs like celeborn test should change to using Spark 3.3 or higher supported versions, instead of deleting those tests.

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks @philo-he Agreed.
Also, created a linked issue for Spark 3.2-specific compatibility code removal: #11379

Comment thread .idea/vcs.xml

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Please confirm if these changes were intended.

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks @philo-he I'm working on it and will commit again.

@QCLyu QCLyu reopened this Jan 16, 2026
@github-actions

Copy link
Copy Markdown

Run Gluten Clickhouse CI on x86

1 similar comment
@github-actions

Copy link
Copy Markdown

Run Gluten Clickhouse CI on x86

@QCLyu
QCLyu marked this pull request as ready for review January 16, 2026 05:51

@philo-he philo-he left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Looks good. Just two minor comments. Thanks.

@zhouyuan, @zzcclp, do you have any comments?

Comment thread .github/workflows/velox_backend_x86.yml Outdated
MVN_CMD: 'build/mvn -ntp'
WGET_CMD: 'wget -nv'
CCACHE_DIR: "${{ github.workspace }}/.ccache"
SETUP: 'source .github/workflows/util/setup-helper.sh'

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Remove this line if not required.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

fixed

| build_velox_tests | Build Velox tests. | OFF |
| build_velox_benchmarks | Build Velox benchmarks (velox_tests and connectors will be disabled if ON) | OFF |
| build_arrow | Build arrow java/cpp and install the libs in local. Can turn it OFF after first build. | ON |
| spark_version | Build for specified version of Spark(3.3, 3.4, 3.5, ALL). `ALL` means build for all versions. | ALL |

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Could you update this line to include 4.0, 4.1?

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Will fix in another PR with doc updates for spark-4.0 update

@zhouyuan zhouyuan left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

👍

@github-actions

Copy link
Copy Markdown

Run Gluten Clickhouse CI on x86

@github-actions

Copy link
Copy Markdown

Run Gluten Clickhouse CI on x86

Signed-off-by: Yuan <yuanzhou@apache.org>
@github-actions

Copy link
Copy Markdown

Run Gluten Clickhouse CI on x86

@zhouyuan
zhouyuan merged commit 7d02d2b into apache:main Jan 20, 2026
106 of 107 checks passed
@zhouyuan

Copy link
Copy Markdown
Member

@QCLyu Thank you for your outstanding work and relentless iteration!

@QCLyu

QCLyu commented Jan 20, 2026

Copy link
Copy Markdown
Contributor Author

Thanks @philo-he @zhouyuan @zzcclp . That's very helpful.

LuciferYang added a commit to LuciferYang/gluten that referenced this pull request Jul 15, 2026
…core

Spark 3.2 support was removed in prior PRs (apache#11351, apache#11687, apache#11731, apache#11887);
currently supported versions are Spark 3.3, 3.4, 3.5, 4.0, 4.1. A few symbols
that existed only as pre-Spark-3.3 shims are still around and dead code today.

1. `GlutenPlan.SupportsRowBasedCompatible` trait

   Introduced to provide `def supportsRowBased(): Boolean` for Spark < 3.3
   where `SparkPlan.supportsRowBased` did not exist yet. The default body was
   `throw new GlutenException("Illegal state: The method is not expected to
   be called")`. On Spark 3.3+, `SparkPlan.supportsRowBased` is native and
   every concrete GlutenPlan / ColumnarInputAdapter overrides it directly, so
   the trait's default was already unreachable. Drop the trait and its two
   mixin sites.

2. `SparkVersionUtil.gteSpark33`

   Always true after Spark 3.2 was dropped. Also drop the single caller
   guard in `canPropagateConvention` (Transitions), which no longer needs to
   skip UnionExec on Spark 3.2. `eqSpark33` and `comparedWithSpark33` are
   kept: they distinguish Spark 3.3 from 3.4+ (different `TaskContextImpl`
   ctor signature and different write planning API), which is unrelated to
   the Spark 3.2 residual concern.

3. `SparkPlanUtil.supportsRowBased` reflection

   The reflection was needed on Spark 3.2 because `SparkPlan.supportsRowBased`
   did not exist as a member yet; the same compiled artifact ran on 3.2 and
   3.3+ only by resolving the method reflectively at call time. Now that
   Spark 3.2 is dropped, a direct call `plan.supportsRowBased` compiles on
   all supported profiles and is strictly better (primitive Boolean instead
   of boxed, no per-call `getMethod` lookup, no `InvocationTargetException`
   wrapping). The 3 callers in `ConventionFunc` are on the planning hot
   path.

Verified via compile on Spark 3.3, 3.5, and 4.1 (scala-2.13) profiles, plus
gluten-core and gluten-substrait tests on Spark 3.5.
LuciferYang added a commit to LuciferYang/gluten that referenced this pull request Jul 15, 2026
Spark 3.2 was dropped by apache#11351/apache#11687/apache#11731/apache#11887; currently supported
versions per README.md / docs/index.md / pom.xml profiles are Spark 3.3,
3.4, 3.5, 4.0, and 4.1. Various docs, comments, and small code paths still
carry Spark 3.2 leftovers or predate the 4.0/4.1 additions. This PR aligns
them.

Docs:
- `docs/get-started/Velox.md`: version table + prose updated to
  3.3.1, 3.4.4, 3.5.5, 4.0.2, 4.1.1 (matching `<spark.version>` in each
  profile).
- `docs/velox-backend-limitations.md`: three section headers "For
  Spark3.2 and Spark3.3" reduced to "For Spark3.3"; the "spark3.2/3.3"
  runtime warning reduced to "spark3.3".
- `docs/get-started/VeloxQAT.md`: replaced hardcoded old version list
  with "all supported Spark versions" since the script iterates
  `SUPPORTED_SPARK_VERSIONS`.
- `docs/developers/HowToRelease.md`: removed the stale `spark-3.2.tar.gz`
  example line.
- `tools/gluten-it/README.md`: profile list updated to
  spark-3.3, spark-3.4, spark-3.5, spark-4.0, spark-4.1.
- `.github/ISSUE_TEMPLATE/bug.yml`: dropdown updated to
  Spark-3.3.x through Spark-4.1.x (removed Spark-3.2.x, added
  Spark-4.1.x, normalized case).

Comments / small code:
- `shims/spark33/.../OrcFileFormat.scala`: comment no longer references
  Spark 3.2 or a hypothetical shims-spark32.
- `backends-velox/.../RowToVeloxColumnarExec.scala`: removed
  `// For spark 3.2.` above `withNewChildInternal`; the override is the
  standard Spark 3.3+ API.
- `backends-velox/.../CudfNodeValidationRule.scala`: replaced
  `.find(_).isDefined` (used because Spark 3.2 lacked `TreeNode.exists`)
  with `.exists(_)`; dropped the stale comment.
- `backends-velox/.../VeloxAggregateFunctionsSuite.scala`: dropped
  stale "Spark 3.2 does not have this configuration" comment.
- `backends-velox/.../ArithmeticAnsiValidateSuite.scala`: comment
  narrowed from "Spark 3.2 and 3.3" to "Spark 3.3".
- `tools/gluten-it/common/.../SparkJvmOptions.java`: dropped the
  Spark-3.2 `ClassNotFoundException` fallback that returned "";
  consolidated all reflection exceptions into a single multi-catch.
LuciferYang added a commit to LuciferYang/gluten that referenced this pull request Jul 15, 2026
Spark 3.2 was dropped by apache#11351/apache#11687/apache#11731/apache#11887; currently supported
versions per README.md / docs/index.md / pom.xml profiles are Spark 3.3,
3.4, 3.5, 4.0, and 4.1. Various docs, comments, and small code paths still
carry Spark 3.2 leftovers or predate the 4.0/4.1 additions. This PR aligns
them.

Docs:
- `docs/get-started/Velox.md`: version table + prose updated to
  3.3.1, 3.4.4, 3.5.5, 4.0.2, 4.1.1 (matching `<spark.version>` in each
  profile).
- `docs/velox-backend-limitations.md`: three section headers "For
  Spark3.2 and Spark3.3" reduced to "For Spark3.3"; the "spark3.2/3.3"
  runtime warning reduced to "spark3.3".
- `docs/get-started/VeloxQAT.md`: replaced hardcoded old version list
  with "all supported Spark versions" since the script iterates
  `SUPPORTED_SPARK_VERSIONS`.
- `docs/developers/HowToRelease.md`: removed the stale `spark-3.2.tar.gz`
  example line.
- `tools/gluten-it/README.md`: profile list updated to
  spark-3.3, spark-3.4, spark-3.5, spark-4.0, spark-4.1.
- `.github/ISSUE_TEMPLATE/bug.yml`: dropdown updated to
  Spark-3.3.x through Spark-4.1.x (removed Spark-3.2.x, added
  Spark-4.1.x, normalized case).

Comments / small code:
- `shims/spark33/.../OrcFileFormat.scala`: comment no longer references
  Spark 3.2 or a hypothetical shims-spark32.
- `backends-velox/.../RowToVeloxColumnarExec.scala`: removed
  `// For spark 3.2.` above `withNewChildInternal`; the override is the
  standard Spark 3.3+ API.
- `backends-velox/.../CudfNodeValidationRule.scala`: replaced
  `.find(_).isDefined` (used because Spark 3.2 lacked `TreeNode.exists`)
  with `.exists(_)`; dropped the stale comment.
- `backends-velox/.../VeloxAggregateFunctionsSuite.scala`: dropped
  stale "Spark 3.2 does not have this configuration" comment.
- `backends-velox/.../ArithmeticAnsiValidateSuite.scala`: comment
  narrowed from "Spark 3.2 and 3.3" to "Spark 3.3".
- `tools/gluten-it/common/.../SparkJvmOptions.java`: dropped the
  Spark-3.2 `ClassNotFoundException` fallback that returned "";
  consolidated all reflection exceptions into a single multi-catch.
zzcclp pushed a commit that referenced this pull request Jul 16, 2026
…12524)

[MINOR][CH] Remove orphan Spark 3.2 / Delta 2.0 source directories (#12524)

The `<delta.binary.version>20</delta.binary.version>` property lived only
in the `spark-3.2` Maven profile, which was removed by #11351. Since then,
`src-delta20/` under both `backends-clickhouse/` and `gluten-delta/` has
been unreachable dead code:

- No profile currently sets `delta.binary.version=20`
  (spark-3.3 -> 23, spark-3.4 -> 24, spark-3.5 -> 33,
   spark-4.0/4.1 -> 40).
- No CI job, doc, or script activates the directories.
- Every deleted class has a live replacement in `src-delta23/` and
  `src-delta33/` (and, for `gluten-delta`, `src-delta24/` /
  `src-delta40/` as well), so no user-visible symbol is lost.

Verified with a clean rebuild on spark-3.3 + backends-clickhouse + delta.
LuciferYang added a commit to LuciferYang/gluten that referenced this pull request Jul 16, 2026
Spark 3.2 was dropped by apache#11351 / apache#11687 / apache#11731 / apache#11887; the
`spark32` protected def in `GlutenClickHouseWholeStageTransformerSuite`
was defined as `sparkVersion.equals("3.2")` and has been dead since.
This PR removes the definition and prunes every reachable
`if (spark32) ...` / `if (!spark32) ...` / `${if (spark32) ... else ...}`
branch to keep only the Spark 3.3+ path.

Test-code changes (all in `backends-clickhouse/src/test/`):

- `GlutenClickHouseWholeStageTransformerSuite.scala`: drop the
  `protected def spark32` definition (always false).
- `GlutenClickHouseTPCHBucketSuite.scala`: `hasSortByCol = !spark32`
  collapses to `true` on all supported Sparks; version-gated
  `if (spark32) ...` branches removed.
- `GlutenClickHouseTPCHParquetBucketSuite.scala`,
  `GlutenClickHouseDeltaParquetWriteSuite.scala`,
  `GlutenClickHouseMergeTreeWriteSuite.scala`,
  `GlutenClickHouseMergeTreeOptimizeSuite.scala`,
  `GlutenClickHouseMergeTreeWriteOnHDFSSuite.scala`,
  `GlutenClickHouseMergeTreeWriteOnHDFSWithRocksDBMetaSuite.scala`,
  `GlutenClickHouseMergeTreeWriteOnS3Suite.scala`,
  `GlutenClickHouseMergeTreePathBasedWriteSuite.scala`: every reachable
  `if (spark32) ... else ...` block, `if (!spark32) ...` guard, and
  inline `${if (spark32) "" else "SORTED BY (...)"}` interpolation is
  reduced to the Spark 3.3+ path (always emit `SORTED BY`).
- `GlutenClickHouseTPCDSParquetAQESuite.scala`,
  `GlutenClickHouseTPCDSParquetColumnarShuffleAQESuite.scala`:
  comments narrowed from "On Spark 3.2, ... on Spark 3.3, ..." to
  describe only the surviving Spark 3.3+ shape.
- `hive/GlutenClickHouseNativeWriteTableSuite.scala`: drop stale
  `// spark 3.2 without orc or parquet suffix` comment.

Main-code change (one file):

- `RowToCHNativeColumnarExec.scala`: drop the `// For spark 3.2.`
  comment above `withNewChildInternal`. The override is required by
  `TreeNode`'s API on every Spark version currently supported by
  Gluten, not a Spark 3.2-only quirk. Mirrors the same cleanup for
  `RowToVeloxColumnarExec` included in apache#12525.

Explicitly kept for a separate follow-up PR:

- `backends-clickhouse/.../ExtendedColumnPruning.scala:66-72` — the
  local `getAttributeToExtractValues` re-implementation exists because
  Spark 3.2's upstream signature was 2-arg. On 3.3+ it is 3-arg; the
  local copy could be replaced with a delegate. That is a real
  refactor, not a comment fix.
- `backends-clickhouse/.../CHColumnarWrite.scala:157` — the
  `bucketSpec` reflection was needed for Spark 3.2, may be replaceable
  with direct access on 3.3+. Also a real refactor.
- `CustomSum.scala:28` — historical provenance of a copied file, not
  a version gate; keep as-is.

Verified with `mvn -pl backends-clickhouse -am install -Pspark-3.3,
backends-clickhouse,delta` (SUCCESS) and `scalastyle:check spotless:check`
(SUCCESS).
LuciferYang added a commit to LuciferYang/gluten that referenced this pull request Jul 16, 2026
Spark 3.2 was dropped by apache#11351 / apache#11687 / apache#11731 / apache#11887; the
`spark32` protected def in `GlutenClickHouseWholeStageTransformerSuite`
was defined as `sparkVersion.equals("3.2")` and has been dead since.
This PR removes the definition and prunes every reachable
`if (spark32) ...` / `if (!spark32) ...` / `${if (spark32) ... else ...}`
branch to keep only the Spark 3.3+ path.

Test-code changes (all in `backends-clickhouse/src/test/`):

- `GlutenClickHouseWholeStageTransformerSuite.scala`: drop the
  `protected def spark32` definition (always false).
- `GlutenClickHouseTPCHBucketSuite.scala`: `hasSortByCol = !spark32`
  collapses to `true` on all supported Sparks; version-gated
  `if (spark32) ...` branches removed.
- `GlutenClickHouseTPCHParquetBucketSuite.scala`,
  `GlutenClickHouseDeltaParquetWriteSuite.scala`,
  `GlutenClickHouseMergeTreeWriteSuite.scala`,
  `GlutenClickHouseMergeTreeOptimizeSuite.scala`,
  `GlutenClickHouseMergeTreeWriteOnHDFSSuite.scala`,
  `GlutenClickHouseMergeTreeWriteOnHDFSWithRocksDBMetaSuite.scala`,
  `GlutenClickHouseMergeTreeWriteOnS3Suite.scala`,
  `GlutenClickHouseMergeTreePathBasedWriteSuite.scala`: every reachable
  `if (spark32) ... else ...` block, `if (!spark32) ...` guard, and
  inline `${if (spark32) "" else "SORTED BY (...)"}` interpolation is
  reduced to the Spark 3.3+ path (always emit `SORTED BY`).
- `GlutenClickHouseTPCDSParquetAQESuite.scala`,
  `GlutenClickHouseTPCDSParquetColumnarShuffleAQESuite.scala`:
  comments narrowed from "On Spark 3.2, ... on Spark 3.3, ..." to
  describe only the surviving Spark 3.3+ shape.
- `hive/GlutenClickHouseNativeWriteTableSuite.scala`: drop stale
  `// spark 3.2 without orc or parquet suffix` comment.

Main-code change (one file):

- `RowToCHNativeColumnarExec.scala`: drop the `// For spark 3.2.`
  comment above `withNewChildInternal`. The override is required by
  `TreeNode`'s API on every Spark version currently supported by
  Gluten, not a Spark 3.2-only quirk. Mirrors the same cleanup for
  `RowToVeloxColumnarExec` included in apache#12525.

Explicitly kept for a separate follow-up PR:

- `backends-clickhouse/.../ExtendedColumnPruning.scala:66-72` — the
  local `getAttributeToExtractValues` re-implementation exists because
  Spark 3.2's upstream signature was 2-arg. On 3.3+ it is 3-arg; the
  local copy could be replaced with a delegate. That is a real
  refactor, not a comment fix.
- `backends-clickhouse/.../CHColumnarWrite.scala:157` — the
  `bucketSpec` reflection was needed for Spark 3.2, may be replaceable
  with direct access on 3.3+. Also a real refactor.
- `CustomSum.scala:28` — historical provenance of a copied file, not
  a version gate; keep as-is.

Verified with `mvn -pl backends-clickhouse -am install -Pspark-3.3,
backends-clickhouse,delta` (SUCCESS) and `scalastyle:check spotless:check`
(SUCCESS).
zzcclp pushed a commit that referenced this pull request Jul 16, 2026
…12532)

[MINOR][CH] Remove residual Spark 3.2 branches from clickhouse tests (#12532)

Spark 3.2 was dropped by #11351 / #11687 / #11731 / #11887; the
`spark32` protected def in `GlutenClickHouseWholeStageTransformerSuite`
was defined as `sparkVersion.equals("3.2")` and has been dead since.
This PR removes the definition and prunes every reachable
`if (spark32) ...` / `if (!spark32) ...` / `${if (spark32) ... else ...}`
branch to keep only the Spark 3.3+ path.

Test-code changes (all in `backends-clickhouse/src/test/`):

- `GlutenClickHouseWholeStageTransformerSuite.scala`: drop the
  `protected def spark32` definition (always false).
- `GlutenClickHouseTPCHBucketSuite.scala`: `hasSortByCol = !spark32`
  collapses to `true` on all supported Sparks; version-gated
  `if (spark32) ...` branches removed.
- `GlutenClickHouseTPCHParquetBucketSuite.scala`,
  `GlutenClickHouseDeltaParquetWriteSuite.scala`,
  `GlutenClickHouseMergeTreeWriteSuite.scala`,
  `GlutenClickHouseMergeTreeOptimizeSuite.scala`,
  `GlutenClickHouseMergeTreeWriteOnHDFSSuite.scala`,
  `GlutenClickHouseMergeTreeWriteOnHDFSWithRocksDBMetaSuite.scala`,
  `GlutenClickHouseMergeTreeWriteOnS3Suite.scala`,
  `GlutenClickHouseMergeTreePathBasedWriteSuite.scala`: every reachable
  `if (spark32) ... else ...` block, `if (!spark32) ...` guard, and
  inline `${if (spark32) "" else "SORTED BY (...)"}` interpolation is
  reduced to the Spark 3.3+ path (always emit `SORTED BY`).
- `GlutenClickHouseTPCDSParquetAQESuite.scala`,
  `GlutenClickHouseTPCDSParquetColumnarShuffleAQESuite.scala`:
  comments narrowed from "On Spark 3.2, ... on Spark 3.3, ..." to
  describe only the surviving Spark 3.3+ shape.
- `hive/GlutenClickHouseNativeWriteTableSuite.scala`: drop stale
  `// spark 3.2 without orc or parquet suffix` comment.

Main-code change (one file):

- `RowToCHNativeColumnarExec.scala`: drop the `// For spark 3.2.`
  comment above `withNewChildInternal`. The override is required by
  `TreeNode`'s API on every Spark version currently supported by
  Gluten, not a Spark 3.2-only quirk. Mirrors the same cleanup for
  `RowToVeloxColumnarExec` included in #12525.

Explicitly kept for a separate follow-up PR:

- `backends-clickhouse/.../ExtendedColumnPruning.scala:66-72` — the
  local `getAttributeToExtractValues` re-implementation exists because
  Spark 3.2's upstream signature was 2-arg. On 3.3+ it is 3-arg; the
  local copy could be replaced with a delegate. That is a real
  refactor, not a comment fix.
- `backends-clickhouse/.../CHColumnarWrite.scala:157` — the
  `bucketSpec` reflection was needed for Spark 3.2, may be replaceable
  with direct access on 3.3+. Also a real refactor.
- `CustomSum.scala:28` — historical provenance of a copied file, not
  a version gate; keep as-is.

Verified with `mvn -pl backends-clickhouse -am install -Pspark-3.3,
backends-clickhouse,delta` (SUCCESS) and `scalastyle:check spotless:check`
(SUCCESS).
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

[VL] deprecated Spark-3.2 unit tests

4 participants