Repository navigation
[VL] Spark 4.2: align profile dependency versions (Hadoop 3.5.0, Jackson 2.21.2, Guava, log4j, commons-lang3) - #13161
Merged
Conversation
akshaytayal
force-pushed
the
spark42-dep-upgrades-main
branch
from
September 29, 2026 15:17
c894982 to
6c64a73
Compare
akshaytayal
force-pushed
the
spark42-dep-upgrades-main
branch
from
September 29, 2026 15:21
6c64a73 to
88da8b4
Compare
Align every Spark-provided library the Gluten spark-4.2 build resolves with the
versions Spark 4.2.0 ships (verified against apache/spark v4.2.0 pom.xml and via
mvn help:evaluate -Pspark-4.2,scala-2.13), plus fix the compile break the Hadoop
bump introduces.
Effective versions for a spark-4.2 build, now all matching Spark 4.2.0:
- Scala 2.13.17 -> 2.13.18 (scoped to spark-4.2 profile; overrides the
shared scala-2.13 profile only for 4.2, other Spark versions keep 2.13.17)
- Hadoop 3.4.1 -> 3.5.0
- Jackson 2.18.2 -> 2.21.2 (annotations pinned to 2.21; no 2.21.2 patch exists)
- slf4j 2.0.16 -> 2.0.17
- Guava 33.4.0-jre -> 33.6.0-jre
- log4j 2.24.3 -> 2.25.4
- commons-lang3 3.17.0 -> 3.20.0
- antlr4 4.13.1, arrow 19.0.0: already matched, unchanged.
jackson-annotations has no 2.21.2 patch release, so it is pinned separately to
2.21 via a new fasterxml.annotations.version property (defaults to
${fasterxml.version} globally -> no-op for other profiles).
VeloxSparkPlanExecApi.scala: Hadoop 3.5.0 drops javax.ws.rs from the compile
classpath, so the spill-path URI rewrite switches from javax.ws.rs.core.UriBuilder
to java.net.URI (JDK-native, same encoded result).
Delta/Iceberg/Hudi are external integrations (not Spark-bundled and not in Spark's
release), so they are intentionally left unchanged.
akshaytayal
force-pushed
the
spark42-dep-upgrades-main
branch
from
September 29, 2026 15:36
88da8b4 to
de279fc
Compare
|
Run Gluten Clickhouse CI on x86 |
|
Run Gluten Clickhouse CI on x86 |
2 similar comments
|
Run Gluten Clickhouse CI on x86 |
|
Run Gluten Clickhouse CI on x86 |
This was referenced Sep 30, 2026
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
What changes were proposed in this pull request?
Aligns every Spark-provided library the Gluten
spark-4.2build resolves with the versionsSpark 4.2.0 ships, and fixes one compile break the Hadoop bump introduces.
Spark 4.2 Lib upgrade release notes: https://spark.apache.org/releases/spark-release-4-2-0.html
Versions were taken from
apache/sparkv4.2.0pom.xml(the release tag)and cross-checked with
mvn help:evaluate -Pspark-4.2,scala-2.13, so theeffective build versions match — not just what a property literal says.
Scoping notes:
scala-2.13profile (2.13.17). It is overriddento
2.13.18inside thespark-4.2profile, so only Spark 4.2 builds pickit up — Spark 3.4/3.5/4.0/4.1 keep 2.13.17 (verified via
help:evaluate).2.21.2patch release (published only at theminor level,
2.21), so it is pinned separately to2.21via a newfasterxml.annotations.versionproperty that defaults to${fasterxml.version}at the top level — a no-op for other profiles.
Spark's release notes), so they are intentionally left unchanged.
VeloxSparkPlanExecApi.scala— jakarta-safe spill path: Hadoop 3.5.0 dropsjavax.ws.rsfrom the compile classpath, so the spill-path URI rewrite switchesfrom
javax.ws.rs.core.UriBuildertojava.net.URI(JDK-native, same encodedresult).
Why are the changes needed?
The Spark 4.2 profile added in #13126 reused the 4.1 dependency versions. Spark
4.2.0 upgrades these libraries; aligning the profile keeps Gluten's Spark 4.2
build consistent with the Spark runtime it targets and avoids classpath/version
skew.
How was this patch tested?
Full CI on a fork (spark-test 3.4/3.5/4.0/4.1 + scala-2.13 + slow/extended,
tpc-test ubuntu/centos). Only the
spark-4.2profile is changed (plus theannotations property default, a no-op for other profiles); effective versions
for other Spark builds are unchanged (confirmed with
mvn help:evaluate).