Skip to content

[CORE] Drop 15.0.0-gluten Arrow version rename and depend on vanilla Apache Arrow - #12244

Merged
philo-he merged 4 commits into
apache:mainfrom
sezruby:arrow-drop-gluten-rename
Jun 9, 2026
Merged

philo-he merged 4 commits into
apache:mainfrom
sezruby:arrow-drop-gluten-rename

Conversation

@sezruby

@sezruby sezruby commented Jun 5, 2026 •

Copy link
Copy Markdown
Contributor

What changes were proposed in this pull request?

Make gluten depend on vanilla Apache Arrow from Maven Central instead of the locally-built 15.0.0-gluten artifact, so x86_64 / aarch64 contributors don't need to run dev/build-arrow.sh's Java build.

Three commits:

  1. Drop the 15.0.0-gluten rename. Removes the arrow-gluten.version property from the parent / Spark-4.x profile poms; switches every gluten-arrow/pom.xml Arrow dep to the vanilla ${arrow.version} coordinate (15.0.0 for Spark 3.x default; 18.1.0 already in the Spark-4.0/4.1 profiles); drops the versions:set -DnewVersion=15.0.0-gluten step in build-arrow.sh. (Note: [MINOR][VL] Drop modify_arrow_dataset_scan_option.patch #12148 already removed the modify_arrow_dataset_scan_option.patch upstream — picked up via merge.)
  2. Skip local Arrow Java build on x86_64 / aarch64. build_arrow_java() now early-returns on non-ppc64le. Maven Central's arrow-c-data:15.0.0 already ships libarrow_cdata_jni for x86_64/ (Linux/macOS/Windows) and aarch_64/ (Linux/macOS) — no local build needed. Saves ~10 min per dev bootstrap. ppc64le path is byte-for-byte unchanged (Central's jar has no ppcle_64/ native, so support_ibm_power.patch + the local mvn install are still required there).
  3. Stop sharing pre-built Arrow Java jars between CI lanes. Removes the mkdir + cp lines that staged Arrow jars, the arrow-jars-* upload-artifact steps, and every downstream Download Arrow Jars step (~30 occurrences across 4 PR-CI workflows). Test lanes now resolve Arrow from Maven Central, which makes CI green a real signal that the change works against the published artifacts. Also cleaned up a stale ls -l .../arrow-dataset/15.0.0-gluten/ debug command.

Patch audit (for the artifact rename drop)

build-arrow.sh previously applied four patches and renamed the resulting jars to 15.0.0-gluten:

Patch Lines Touches Status
modify_arrow.patch 135 C++ only Still applied
modify_arrow_dataset_scan_option.patch 883 Adds dead Java classes; C++ Substrait Removed by #12148 upstream
cmake-compatibility.patch 34 C++ only Unchanged
support_ibm_power.patch 28 ppc64le JniLoader switch case Still applied on ppc64le; doesn't need a custom artifact coordinate

A sweep (grep -rn 'org.apache.arrow.dataset' --include='*.java' --include='*.scala') confirms the only main-source consumers of arrow-dataset are gluten-arrow/.../ArrowNativeMemoryPool and ArrowReservationListener, which use only the upstream org.apache.arrow.dataset.jni.{NativeMemoryPool, ReservationListener} types — no patched classes.

Effect on contributors

  • x86_64 / aarch64: dev/build-arrow.sh is no longer required for the JVM Arrow side. Local builds skip the ~10-min Arrow Java compile.
  • ppc64le: dev/build-arrow.sh's C++ build still runs (Velox links static cpp Arrow); its Java build still runs (only on ppc64le, gated by uname -m). Same dev loop as today.

Effect on shading

Independent of bundling. The bundled gluten-velox-bundle still ships unshaded org.apache.arrow.* per #12226. Unbundling Arrow entirely (relying on Spark's shipped Arrow at runtime) was tried in #12245 — closed because Spark 3.3 / 3.4 ship Arrow 7 / 11 and the LCD-pin trade-offs aren't worth it. Re-evaluate when those versions are dropped.

How was this patch tested?

  • mvn dependency:tree -pl gluten-arrow shows every org.apache.arrow:* resolved at vanilla 15.0.0 / 18.1.0 from Central.
  • Sweep for stale references: grep -rn 'arrow-gluten\.version\|15\.0\.0-gluten' — no matches.
  • CI: with commit 3, every test lane now downloads Arrow from Maven Central rather than from the workspace artifact, so a green run validates the Central-only path end-to-end.

References

… Apache Arrow

The custom 15.0.0-gluten artifact coordinate forced every contributor to run
dev/build-arrow.sh before they could build gluten, even though the Java side
of that build no longer carries any load-bearing modifications:

* The 883-line modify_arrow_dataset_scan_option.patch added CSV / Substrait
  dataset Java classes (CsvFragmentScanOptions, ConvertUtil, etc.). Every
  consumer of those classes inside gluten was deleted by apache#12130 along with
  the Arrow-CSV / Arrow-Dataset JVM code path. The patch is no longer applied
  to the Arrow Java build here; the file itself is kept because get-velox.sh
  still copies it into Velox's CMake Arrow EP for the C++ side.
* support_ibm_power.patch (ppc64le → ppcle_64 in JniLoader) is still load
  bearing for ppc64le builds, but does not require an artifact rename — it
  only patches the binary-resource lookup inside the arrow-c-data JNI jar
  and is still applied by build-arrow.sh.
* The C++ patches (modify_arrow.patch, cmake-compatibility.patch) are
  unchanged.

After this change, on x86_64 / aarch64 every gluten-arrow Arrow dependency
resolves from Maven Central (arrow-c-data:15.0.0, arrow-dataset:15.0.0,
arrow-vector:15.0.0, arrow-memory-{core,unsafe,netty}:15.0.0; 18.1.0 for
the Spark 4.x profiles). ppc64le builds still rely on dev/build-arrow.sh
to produce locally-patched 15.0.0 artifacts — the local-m2 install
overrides Central as before.

Note: this PR removes the artifact-rename indirection but does not yet
unbundle Arrow from the gluten-velox bundle. The bundle still ships
unshaded Arrow (per apache#12226) at the same vanilla coordinates. Removing
the bundled Arrow in favour of Spark's bundled copy is a separate
follow-up driven by the discussion on apache#12226.
@github-actions github-actions Bot added CORE works for Gluten Core BUILD VELOX labels Jun 5, 2026
@github-actions

github-actions Bot commented Jun 5, 2026

Copy link
Copy Markdown

Run Gluten Clickhouse CI on x86

Comment thread dev/build-arrow.sh
Comment on lines 113 to 116
# Arrow Java libraries
${MVN_CMD} install -Parrow-jni -P arrow-c-data -pl c,dataset -am \
-Darrow.c.jni.dist.dir=$ARROW_INSTALL_DIR/lib -Darrow.dataset.jni.dist.dir=$ARROW_INSTALL_DIR/lib -Darrow.cpp.build.dir=$ARROW_INSTALL_DIR/lib \
-Dmaven.test.skip -Drat.skip -Dmaven.gitcommitid.skip -Dcheckstyle.skip -Dassembly.skipAssembly

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Do we still need to build Arrow Java locally?

@sezruby

sezruby commented Jun 5, 2026

Copy link
Copy Markdown
Contributor Author

Do we still need to build Arrow Java locally?

Mostly no — Maven Central's arrow-c-data:15.0.0 jar already ships libarrow_cdata_jni for x86_64/ (Linux/macOS/Windows) and aarch_64/ (Linux/macOS), so x86_64 / aarch64 contributors no longer need the local Java build after this PR.

The reason it's still wired into dev/builddeps-veloxbe.sh unconditionally:

  • ppc64le has no native in the Central jar. support_ibm_power.patch (kept) adds the ppc64le → ppcle_64 arch case to JniLoader.java and the local mvn install step bakes a locally-built libarrow_cdata_jni.so for ppc64le into the resulting arrow-c-data:15.0.0 jar in ~/.m2, overriding Central.

Happy to add a follow-up commit gating build_arrow_java on [[ $(uname -m) == ppc64le ]] so x86_64 / aarch64 users skip ~10 min of redundant work. Holding off in this PR because there's no ppc64le CI lane to confirm the conditional doesn't break the patched build, and I don't have a qemu setup locally to validate it either.

@FelixYBW

FelixYBW commented Jun 8, 2026

Copy link
Copy Markdown
Contributor
  • ppc64le has no native in the Central jar.

@Jenkins-J Can you fix this?

@FelixYBW

FelixYBW commented Jun 8, 2026

Copy link
Copy Markdown
Contributor

Spark 3.3.1 ships Arrow 7.0.0
Spark 3.4.4 ships Arrow 11.0.0
Spark 3.5.5 ships Arrow 15.0.0
Spark 4.0 / 4.1 ship Arrow 18.x

Per @zhztheplayer, once we completely remove the csv reader, we should be able to use Arrow 7 and Arrow 11.

I re-open the PR: #12148

@zhztheplayer

Copy link
Copy Markdown
Member

Happy to add a follow-up commit gating build_arrow_java on [[ $(uname -m) == ppc64le ]] so x86_64 / aarch64 users skip ~10 min of redundant work.

@sezruby we can give this a try, thanks.

We'd also keep an eye on the glibc compatibility of the official arrow-c-data jar.

Maven Central's arrow-c-data / arrow-dataset jars at the pinned version already
ship libarrow_cdata_jni and libarrow_dataset_jni for x86_64 (Linux / macOS /
Windows) and aarch_64 (Linux / macOS), so contributors on those archs no longer
need build-arrow.sh's mvn install step — gluten-arrow resolves the same
artifact transitively from Central.

Skip build_arrow_java when uname -m is not ppc64le; ppc64le still needs the
local install because Central's jar carries no ppcle_64 native, and
support_ibm_power.patch (kept) adds that arch case to JniLoader.java.

build_arrow_cpp and prepare_arrow_build stay unconditional — Velox links
against the static C++ Arrow regardless of arch, and the patched source tree
is needed for that build path.

Saves ~10 min of redundant `mvn install` on every dev bootstrap on
x86_64 / aarch64. Behavior on ppc64le is unchanged.

Follow-up to review feedback on the parent commit.
@github-actions

github-actions Bot commented Jun 8, 2026

Copy link
Copy Markdown

Run Gluten Clickhouse CI on x86

@sezruby

sezruby commented Jun 8, 2026

Copy link
Copy Markdown
Contributor Author

@zhztheplayer @philo-he Pushed a follow-up commit (a9d80cce3) that early-returns from build_arrow_java() when uname -m != ppc64le, so x86_64 / aarch64 contributors skip the redundant local install — gluten-arrow resolves arrow-c-data / arrow-dataset from Maven Central on those archs. ppc64le path is byte-for-byte unchanged.

Two CI failures on this run, both unrelated to this PR's diff:

  • spark-test-spark35-slow (2m31s) — failed at container init: Curl error (7): Couldn't connect to server for http://vault.centos.org/centos/8/AppStream/x86_64/os/repodata/repomd.xml. CentOS 8 mirror outage.
  • spark-test-spark41 (22m10s) — failed in a SparkScriptTransformationExec test asserting some_non_existent_command produces a SparkException. The error stack is from Spark's subprocess machinery, not gluten code; looks like an environment / flake.

Could you re-trigger those two lanes when you get a chance?

One caveat worth flagging: every gluten CI lane runs builddeps-veloxbe.sh with BUILD_ARROW=OFF and uses the pre-built Arrow baked into the Docker image, so the BUILD_ARROW=ON path (where build-arrow.sh actually executes) is not covered by CI. That means the conditional I added isn't exercised by any lane — it relies on a safe-by-construction argument: early-return on non-ppc64le; ppc64le branch byte-for-byte unchanged from before. Worth keeping in mind for review, and probably worth a separate followup to add at least one CI lane that does run BUILD_ARROW=ON end-to-end.

@github-actions

github-actions Bot commented Jun 9, 2026

Copy link
Copy Markdown

Run Gluten Clickhouse CI on x86

@philo-he

philo-he commented Jun 9, 2026

Copy link
Copy Markdown
Member

@sezruby, thank you for the continued efforts.

In the CI workflow, should we remove the Arrow artifact sharing from the build job? Including, but not limited to, the following:

cp -r /root/.m2/repository/org/apache/arrow/* /work/.m2/repository/org/apache/arrow/

- uses: actions/upload-artifact@v4
with:
name: arrow-jars-centos-7-${{github.sha}}
path: .m2/repository/org/apache/arrow/
if-no-files-found: error

@philo-he

philo-he commented Jun 9, 2026

Copy link
Copy Markdown
Member

We'd also keep an eye on the glibc compatibility of the official arrow-c-data jar.

@zhztheplayer, based on the following build guide, it appears that java-jni-manylinux-2014 is used in Arrow's official build. Since it is CentOS 7-based and built against glibc 2.17, this should ensure glibc compatibility on higher-version environments.

https://arrow.apache.org/java/main/developers/building.html

@FelixYBW

FelixYBW commented Jun 9, 2026 •

Copy link
Copy Markdown
Contributor

We'd also keep an eye on the glibc compatibility of the official arrow-c-data jar.

@zhztheplayer, based on the following build guide, it appears that java-jni-manylinux-2014 is used in Arrow's official build. Since it is CentOS 7-based and built against glibc 2.17, this should ensure glibc compatibility on higher-version environments.

https://arrow.apache.org/java/main/developers/building.html

In Gluten, when we convert velox to Arrow C_data, we should link the one in Velox, not the lib in arrow jar, right? Then we copy the data pointers to arrow jar's and send to JVM.

If any project uses Spark's Arrow C_data, then we should be fine.

Spark uses 4 arrow libraries: arrow-format-12.0.1.jar arrow-memory-core-12.0.1.jar arrow-memory-netty-12.0.1.jar arrow-vector-12.0.1.jar

c_data lib is in arrow-c-data-12.0.1.jar. If there is libc conflict, we can build the jar only.

@FelixYBW

FelixYBW commented Jun 9, 2026

Copy link
Copy Markdown
Contributor

@sezruby Does lance-spark use arrow-c-data.jar? Did you build locally or download from maven?

@zhztheplayer

Copy link
Copy Markdown
Member

@sezruby Does lance-spark use arrow-c-data.jar?

Yes it does use arrow-c-data https://github.com/lance-format/lance-spark/blob/48ddbd12d5bf28c5d886ed99b0f21a3824ef98b3/pom.xml#L182-L186.

@zhztheplayer

Copy link
Copy Markdown
Member

based on the following build guide, it appears that java-jni-manylinux-2014 is used in Arrow's official build.

Looks good, this should avoid the libc incompatibilities.

@zhztheplayer

zhztheplayer commented Jun 9, 2026 •

Copy link
Copy Markdown
Member

@sezruby

One caveat worth flagging: every gluten CI lane runs builddeps-veloxbe.sh with BUILD_ARROW=OFF and uses the pre-built Arrow baked into the Docker image, so the BUILD_ARROW=ON path (where build-arrow.sh actually executes) is not covered by CI.

Yes, and I think we still need to verify the change on CI. Can you push a debug commit to remove the pre-built Arrow Java jar from local Maven repo, then see if CI can pass?

Once verified, that debug commit can be reverted before merging.

@github-actions

github-actions Bot commented Jun 9, 2026

Copy link
Copy Markdown

Run Gluten Clickhouse CI on x86

@zhztheplayer

zhztheplayer commented Jun 9, 2026 •

Copy link
Copy Markdown
Member

@sezruby

I see @philo-he mentioned there are CI steps copying arrow jars for subsequent CI tests. Can we clean them up in this PR? By doing so I think we don't need a debug commit that I mentioned in my last comment.

By all means we'd verify the official Arrow Java jars on CI.

Previously the build-native lanes copied /root/.m2/repository/org/apache/arrow/*
(pre-built Arrow Java jars baked into the Docker image) into a workspace
.m2 path and uploaded them as an `arrow-jars-*-${sha}` artifact. Every
downstream test lane downloaded that artifact into its own /root/.m2 so its
Maven build resolved Arrow from the pre-built copy instead of Maven Central.

This pre-bake has two problems now:

1. It hides whether gluten can actually resolve Arrow from Central (the
   change in the parent commits depends on this — gluten-arrow now resolves
   arrow-c-data:15.0.0 / arrow-dataset:15.0.0 etc. from Central on
   x86_64 / aarch64).
2. It exercises the locally-patched 15.0.0-gluten Arrow that no longer
   exists after the parent commits remove that artifact rename.

Removed:
- The mkdir + cp lines that staged Arrow jars after each native build
  (velox_backend_x86.yml, velox_backend_arm.yml, velox_backend_ansi.yml,
  velox_backend_enhanced.yml).
- The `arrow-jars-*` upload-artifact steps that published the staged copy.
- Every Download Arrow Jars / Download All Arrow Jar Artifacts step in
  every downstream test lane (~30 occurrences across the four PR-CI
  workflows).
- A leftover `ls -l .../arrow-dataset/15.0.0-gluten/` debug command in
  velox_backend_x86.yml that referenced the now-removed coordinate.

Untouched: velox_nightly.yml and build_bundle_package.yml — those build
release artifacts and may legitimately want to bake-in a specific Arrow.
@sezruby
sezruby force-pushed the arrow-drop-gluten-rename branch from 2a5d395 to 4249dde Compare June 9, 2026 15:28
@github-actions github-actions Bot added the INFRA label Jun 9, 2026
@github-actions

github-actions Bot commented Jun 9, 2026

Copy link
Copy Markdown

Run Gluten Clickhouse CI on x86

@philo-he philo-he changed the title [CORE] Drop 15.0.0-gluten Arrow version rename, depend on vanilla Apache Arrow [CORE] Drop 15.0.0-gluten Arrow version rename and depend on vanilla Apache Arrow Jun 9, 2026

@philo-he philo-he left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks.

@Jenkins-J

Copy link
Copy Markdown
Contributor

@FelixYBW I've created a PR over in the arrow repository related to version 15 (apache/arrow#50138). Let me know if there are any modifications I should make to the PR.

@philo-he
philo-he merged commit 3d585da into apache:main Jun 9, 2026
66 checks passed
@prestodb-ci

prestodb-ci commented Jun 9, 2026 •

Copy link
Copy Markdown

Failed to get lakehouse/gluten rebase URL from queue item URL http://ci.ibm.prestodb.dev/queue/item/375492/: build not available after 6 attempts

felipepessoto added a commit to felipepessoto/gluten that referenced this pull request Jul 9, 2026
… body

Two readability/cleanup changes to the Delta Spark UT pipeline:

1. Drop redundant Arrow jar handling. Since apache#12244
   (2026-06-09) Gluten depends on vanilla Apache Arrow instead of the
   custom 15.0.0-gluten rename, so the Arrow jars resolve from Maven
   Central. Remove the arrow-jar staging in both native builds, the
   velox-arrow-jars / delta-spark-ut-arrow-jars uploads, the caller's
   arrow_jars_artifact input, and the bundle job's "Download Arrow jars"
   step. The bundle build's Maven cache still holds Arrow across runs and
   resolves it from Central on a miss.

2. Extract the shard test body into run-delta-tests.sh. The "Run Delta
   spark module tests" step embedded ~190 lines of shell (hang watchdog,
   sbt invocation with JVM/heap tuning, cgroup memory forensics, report
   check, known-failures gate). Move that whole body into
   util/delta-spark-ut/run-delta-tests.sh so the workflow step is just an
   `env:` block plus a one-line script call, shrinking delta_spark_ut.yml
   by ~180 lines. The move is faithful: the script body is the previous
   inline block with only the GitHub `${{ }}` expressions replaced by env
   vars (matrix.shard -> SHARD_ID which is already job-level env;
   spark_version/update_baseline/fail_on_fixed added to the step env).
   Verified by a textual diff against the old body and by end-to-end runs
   with a stubbed sbt (baseline-only -> green, regression -> red, seed ->
   green).

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
felipepessoto added a commit to felipepessoto/gluten that referenced this pull request Jul 10, 2026
… body

Two readability/cleanup changes to the Delta Spark UT pipeline:

1. Drop redundant Arrow jar handling. Since apache#12244
   (2026-06-09) Gluten depends on vanilla Apache Arrow instead of the
   custom 15.0.0-gluten rename, so the Arrow jars resolve from Maven
   Central. Remove the arrow-jar staging in both native builds, the
   velox-arrow-jars / delta-spark-ut-arrow-jars uploads, the caller's
   arrow_jars_artifact input, and the bundle job's "Download Arrow jars"
   step. The bundle build's Maven cache still holds Arrow across runs and
   resolves it from Central on a miss.

2. Extract the shard test body into run-delta-tests.sh. The "Run Delta
   spark module tests" step embedded ~190 lines of shell (hang watchdog,
   sbt invocation with JVM/heap tuning, cgroup memory forensics, report
   check, known-failures gate). Move that whole body into
   util/delta-spark-ut/run-delta-tests.sh so the workflow step is just an
   `env:` block plus a one-line script call, shrinking delta_spark_ut.yml
   by ~180 lines. The move is faithful: the script body is the previous
   inline block with only the GitHub `${{ }}` expressions replaced by env
   vars (matrix.shard -> SHARD_ID which is already job-level env;
   spark_version/update_baseline/fail_on_fixed added to the step env).
   Verified by a textual diff against the old body and by end-to-end runs
   with a stubbed sbt (baseline-only -> green, regression -> red, seed ->
   green).

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
felipepessoto added a commit to felipepessoto/gluten that referenced this pull request Jul 10, 2026
… body

Two readability/cleanup changes to the Delta Spark UT pipeline:

1. Drop redundant Arrow jar handling. Since apache#12244
   (2026-06-09) Gluten depends on vanilla Apache Arrow instead of the
   custom 15.0.0-gluten rename, so the Arrow jars resolve from Maven
   Central. Remove the arrow-jar staging in both native builds, the
   velox-arrow-jars / delta-spark-ut-arrow-jars uploads, the caller's
   arrow_jars_artifact input, and the bundle job's "Download Arrow jars"
   step. The bundle build's Maven cache still holds Arrow across runs and
   resolves it from Central on a miss.

2. Extract the shard test body into run-delta-tests.sh. The "Run Delta
   spark module tests" step embedded ~190 lines of shell (hang watchdog,
   sbt invocation with JVM/heap tuning, cgroup memory forensics, report
   check, known-failures gate). Move that whole body into
   util/delta-spark-ut/run-delta-tests.sh so the workflow step is just an
   `env:` block plus a one-line script call, shrinking delta_spark_ut.yml
   by ~180 lines. The move is faithful: the script body is the previous
   inline block with only the GitHub `${{ }}` expressions replaced by env
   vars (matrix.shard -> SHARD_ID which is already job-level env;
   spark_version/update_baseline/fail_on_fixed added to the step env).
   Verified by a textual diff against the old body and by end-to-end runs
   with a stubbed sbt (baseline-only -> green, regression -> red, seed ->
   green).

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
felipepessoto added a commit to felipepessoto/gluten that referenced this pull request Jul 13, 2026
… body

Two readability/cleanup changes to the Delta Spark UT pipeline:

1. Drop redundant Arrow jar handling. Since apache#12244
   (2026-06-09) Gluten depends on vanilla Apache Arrow instead of the
   custom 15.0.0-gluten rename, so the Arrow jars resolve from Maven
   Central. Remove the arrow-jar staging in both native builds, the
   velox-arrow-jars / delta-spark-ut-arrow-jars uploads, the caller's
   arrow_jars_artifact input, and the bundle job's "Download Arrow jars"
   step. The bundle build's Maven cache still holds Arrow across runs and
   resolves it from Central on a miss.

2. Extract the shard test body into run-delta-tests.sh. The "Run Delta
   spark module tests" step embedded ~190 lines of shell (hang watchdog,
   sbt invocation with JVM/heap tuning, cgroup memory forensics, report
   check, known-failures gate). Move that whole body into
   util/delta-spark-ut/run-delta-tests.sh so the workflow step is just an
   `env:` block plus a one-line script call, shrinking delta_spark_ut.yml
   by ~180 lines. The move is faithful: the script body is the previous
   inline block with only the GitHub `${{ }}` expressions replaced by env
   vars (matrix.shard -> SHARD_ID which is already job-level env;
   spark_version/update_baseline/fail_on_fixed added to the step env).
   Verified by a textual diff against the old body and by end-to-end runs
   with a stubbed sbt (baseline-only -> green, regression -> red, seed ->
   green).

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
felipepessoto added a commit to felipepessoto/gluten that referenced this pull request Jul 20, 2026
… body

Two readability/cleanup changes to the Delta Spark UT pipeline:

1. Drop redundant Arrow jar handling. Since apache#12244
   (2026-06-09) Gluten depends on vanilla Apache Arrow instead of the
   custom 15.0.0-gluten rename, so the Arrow jars resolve from Maven
   Central. Remove the arrow-jar staging in both native builds, the
   velox-arrow-jars / delta-spark-ut-arrow-jars uploads, the caller's
   arrow_jars_artifact input, and the bundle job's "Download Arrow jars"
   step. The bundle build's Maven cache still holds Arrow across runs and
   resolves it from Central on a miss.

2. Extract the shard test body into run-delta-tests.sh. The "Run Delta
   spark module tests" step embedded ~190 lines of shell (hang watchdog,
   sbt invocation with JVM/heap tuning, cgroup memory forensics, report
   check, known-failures gate). Move that whole body into
   util/delta-spark-ut/run-delta-tests.sh so the workflow step is just an
   `env:` block plus a one-line script call, shrinking delta_spark_ut.yml
   by ~180 lines. The move is faithful: the script body is the previous
   inline block with only the GitHub `${{ }}` expressions replaced by env
   vars (matrix.shard -> SHARD_ID which is already job-level env;
   spark_version/update_baseline/fail_on_fixed added to the step env).
   Verified by a textual diff against the old body and by end-to-end runs
   with a stubbed sbt (baseline-only -> green, regression -> red, seed ->
   green).

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
felipepessoto added a commit to felipepessoto/gluten that referenced this pull request Jul 20, 2026
… body

Two readability/cleanup changes to the Delta Spark UT pipeline:

1. Drop redundant Arrow jar handling. Since apache#12244
   (2026-06-09) Gluten depends on vanilla Apache Arrow instead of the
   custom 15.0.0-gluten rename, so the Arrow jars resolve from Maven
   Central. Remove the arrow-jar staging in both native builds, the
   velox-arrow-jars / delta-spark-ut-arrow-jars uploads, the caller's
   arrow_jars_artifact input, and the bundle job's "Download Arrow jars"
   step. The bundle build's Maven cache still holds Arrow across runs and
   resolves it from Central on a miss.

2. Extract the shard test body into run-delta-tests.sh. The "Run Delta
   spark module tests" step embedded ~190 lines of shell (hang watchdog,
   sbt invocation with JVM/heap tuning, cgroup memory forensics, report
   check, known-failures gate). Move that whole body into
   util/delta-spark-ut/run-delta-tests.sh so the workflow step is just an
   `env:` block plus a one-line script call, shrinking delta_spark_ut.yml
   by ~180 lines. The move is faithful: the script body is the previous
   inline block with only the GitHub `${{ }}` expressions replaced by env
   vars (matrix.shard -> SHARD_ID which is already job-level env;
   spark_version/update_baseline/fail_on_fixed added to the step env).
   Verified by a textual diff against the old body and by end-to-end runs
   with a stubbed sbt (baseline-only -> green, regression -> red, seed ->
   green).

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
felipepessoto added a commit to felipepessoto/gluten that referenced this pull request Jul 26, 2026
… body

Two readability/cleanup changes to the Delta Spark UT pipeline:

1. Drop redundant Arrow jar handling. Since apache#12244
   (2026-06-09) Gluten depends on vanilla Apache Arrow instead of the
   custom 15.0.0-gluten rename, so the Arrow jars resolve from Maven
   Central. Remove the arrow-jar staging in both native builds, the
   velox-arrow-jars / delta-spark-ut-arrow-jars uploads, the caller's
   arrow_jars_artifact input, and the bundle job's "Download Arrow jars"
   step. The bundle build's Maven cache still holds Arrow across runs and
   resolves it from Central on a miss.

2. Extract the shard test body into run-delta-tests.sh. The "Run Delta
   spark module tests" step embedded ~190 lines of shell (hang watchdog,
   sbt invocation with JVM/heap tuning, cgroup memory forensics, report
   check, known-failures gate). Move that whole body into
   util/delta-spark-ut/run-delta-tests.sh so the workflow step is just an
   `env:` block plus a one-line script call, shrinking delta_spark_ut.yml
   by ~180 lines. The move is faithful: the script body is the previous
   inline block with only the GitHub `${{ }}` expressions replaced by env
   vars (matrix.shard -> SHARD_ID which is already job-level env;
   spark_version/update_baseline/fail_on_fixed added to the step env).
   Verified by a textual diff against the old body and by end-to-end runs
   with a stubbed sbt (baseline-only -> green, regression -> red, seed ->
   green).

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
felipepessoto added a commit to felipepessoto/gluten that referenced this pull request Jul 27, 2026
… body

Two readability/cleanup changes to the Delta Spark UT pipeline:

1. Drop redundant Arrow jar handling. Since apache#12244
   (2026-06-09) Gluten depends on vanilla Apache Arrow instead of the
   custom 15.0.0-gluten rename, so the Arrow jars resolve from Maven
   Central. Remove the arrow-jar staging in both native builds, the
   velox-arrow-jars / delta-spark-ut-arrow-jars uploads, the caller's
   arrow_jars_artifact input, and the bundle job's "Download Arrow jars"
   step. The bundle build's Maven cache still holds Arrow across runs and
   resolves it from Central on a miss.

2. Extract the shard test body into run-delta-tests.sh. The "Run Delta
   spark module tests" step embedded ~190 lines of shell (hang watchdog,
   sbt invocation with JVM/heap tuning, cgroup memory forensics, report
   check, known-failures gate). Move that whole body into
   util/delta-spark-ut/run-delta-tests.sh so the workflow step is just an
   `env:` block plus a one-line script call, shrinking delta_spark_ut.yml
   by ~180 lines. The move is faithful: the script body is the previous
   inline block with only the GitHub `${{ }}` expressions replaced by env
   vars (matrix.shard -> SHARD_ID which is already job-level env;
   spark_version/update_baseline/fail_on_fixed added to the step env).
   Verified by a textual diff against the old body and by end-to-end runs
   with a stubbed sbt (baseline-only -> green, regression -> red, seed ->
   green).

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
felipepessoto added a commit to felipepessoto/gluten that referenced this pull request Jul 28, 2026
… body

Two readability/cleanup changes to the Delta Spark UT pipeline:

1. Drop redundant Arrow jar handling. Since apache#12244
   (2026-06-09) Gluten depends on vanilla Apache Arrow instead of the
   custom 15.0.0-gluten rename, so the Arrow jars resolve from Maven
   Central. Remove the arrow-jar staging in both native builds, the
   velox-arrow-jars / delta-spark-ut-arrow-jars uploads, the caller's
   arrow_jars_artifact input, and the bundle job's "Download Arrow jars"
   step. The bundle build's Maven cache still holds Arrow across runs and
   resolves it from Central on a miss.

2. Extract the shard test body into run-delta-tests.sh. The "Run Delta
   spark module tests" step embedded ~190 lines of shell (hang watchdog,
   sbt invocation with JVM/heap tuning, cgroup memory forensics, report
   check, known-failures gate). Move that whole body into
   util/delta-spark-ut/run-delta-tests.sh so the workflow step is just an
   `env:` block plus a one-line script call, shrinking delta_spark_ut.yml
   by ~180 lines. The move is faithful: the script body is the previous
   inline block with only the GitHub `${{ }}` expressions replaced by env
   vars (matrix.shard -> SHARD_ID which is already job-level env;
   spark_version/update_baseline/fail_on_fixed added to the step env).
   Verified by a textual diff against the old body and by end-to-end runs
   with a stubbed sbt (baseline-only -> green, regression -> red, seed ->
   green).

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
felipepessoto added a commit to felipepessoto/gluten that referenced this pull request Jul 28, 2026
… body

Two readability/cleanup changes to the Delta Spark UT pipeline:

1. Drop redundant Arrow jar handling. Since apache#12244
   (2026-06-09) Gluten depends on vanilla Apache Arrow instead of the
   custom 15.0.0-gluten rename, so the Arrow jars resolve from Maven
   Central. Remove the arrow-jar staging in both native builds, the
   velox-arrow-jars / delta-spark-ut-arrow-jars uploads, the caller's
   arrow_jars_artifact input, and the bundle job's "Download Arrow jars"
   step. The bundle build's Maven cache still holds Arrow across runs and
   resolves it from Central on a miss.

2. Extract the shard test body into run-delta-tests.sh. The "Run Delta
   spark module tests" step embedded ~190 lines of shell (hang watchdog,
   sbt invocation with JVM/heap tuning, cgroup memory forensics, report
   check, known-failures gate). Move that whole body into
   util/delta-spark-ut/run-delta-tests.sh so the workflow step is just an
   `env:` block plus a one-line script call, shrinking delta_spark_ut.yml
   by ~180 lines. The move is faithful: the script body is the previous
   inline block with only the GitHub `${{ }}` expressions replaced by env
   vars (matrix.shard -> SHARD_ID which is already job-level env;
   spark_version/update_baseline/fail_on_fixed added to the step env).
   Verified by a textual diff against the old body and by end-to-end runs
   with a stubbed sbt (baseline-only -> green, regression -> red, seed ->
   green).

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
felipepessoto added a commit to felipepessoto/gluten that referenced this pull request Jul 28, 2026
…es baseline

Run delta-io/delta's `spark` ScalaTest suite against a Gluten Velox bundle in CI
and gate the results against a committed baseline so the many expected Delta-on-
Gluten failures stay manageable and can be fixed incrementally without letting
currently-passing tests silently regress.

What it adds (.github/workflows/util/delta-spark-ut/):
- delta_spark_ut.yml: builds the native lib + Gluten bundle, then runs the Delta
  spark suite sharded by suite into 4 shards x 4 forked test JVMs (~16-way), and
  gates each shard against the baseline.
- compare-test-results.py: the gate. Per shard, regressions (failed not in the
  baseline) fail the build; newly-passing baselined tests are flagged so the
  baseline can be tightened. Also supports seed/aggregate modes.
- known-failures.txt: the committed baseline of expected failures.
- setup-delta.sh: clones Delta, injects the Gluten bundle, patches
  DeltaSQLCommandTest, and force-fails the two DeletionVectorsSuite 2B-row tests
  whose native row-index materialization OOM-kills the runner and hangs the shard.
- README.md: how the pipeline, gating and baseline-refresh work.

The workflow also carries a hang watchdog that thread-dumps and kills a wedged
fork, and tunes the per-fork heap (2G) and off-heap (2G) to fit the ~16G runner.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

Change order of steps

Co-authored-by: Copilot Autofix powered by AI <175728472+Copilot@users.noreply.github.com>

[CI] Cherry-pick Delta FileSourceScanLike test fixes and refresh baseline

Delta's data-skipping, limit-push-down, column-pruning and scan-metric tests
collect file-source scans by matching the concrete `FileSourceScanExec` case
class. Under the Gluten Velox bundle the scan is offloaded to
DeltaScanTransformer, a sibling that implements the same `FileSourceScanLike`
interface but is not FileSourceScanExec, so the match misses and the scan
looks absent. This surfaced as `scala.MatchError: List()` (~56
DataSkipping*/DeltaLimitPushDown* tests), empty generated-column partition
filters (~45 OptimizeGeneratedColumnSuite tests) and broken column-pruning /
scan-metric checks across the Delete, Update, Merge, DeletionVectors and
RowId suites and the TestsStatistics helper.

Gluten copies `partitionFilters` and the other accessors these tests read
verbatim onto the offloaded scan, so results are identical to vanilla -- only
the test's `case` match breaks. Fix it by cherry-picking the two merged
upstream Delta commits that widen these matches to the shared
`FileSourceScanLike` interface (behavior-preserving for vanilla, which also
implements it):

  * delta-io/delta#7104 -- ScanReportHelper.collectScans
  * delta-io/delta#7105 -- the remaining 9 test sources, its follow-up

Both are merged on Delta master but land after the ref this workflow builds
against (v4.2.0), so setup-delta.sh cherry-picks them onto the shallow
checkout. Each fetches the fix commit at depth 2 (commit + parent) so
cherry-pick can compute the parent->fix diff, and uses `cherry-pick -n` so no
committer identity is required. Once the pinned DELTA_REF advances to include
a commit its cherry-pick becomes a clean no-op and that block can be removed.

The cherry-picks run before the DeletionVectorsSuite 2B-row force-fail step:
that step sed-injects fail() into DeletionVectorsSuite.scala, which
delta-io/delta#7105 also edits, and git cherry-pick refuses to apply onto a
working tree with uncommitted changes to a file it touches (exit 128).

Refresh known-failures.txt from run 28299900971 (the delta-spark-aggregate job
output), which ran all 19073 tests across 16 shards: removes 187 now-passing
tests with 0 regressions, 963 -> 776. ~147 come from the fixes above
(DataSkipping*, DeltaLimitPushDown*, OptimizeGeneratedColumnSuite, MergeInto*,
RowIdSuite); the remaining ~40 are other suites that now pass (e.g.
HiveConvertToDeltaSuite, BitmapAggregatorE2ESuite). Verified against the
per-shard ran/failed lists: every baseline entry was observed this run (0
stale), so nothing was dropped due to a crashed or incomplete shard.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

[CI] Reuse velox_backend_x86's native build for Delta Spark UT

Make delta_spark_ut.yml a reusable workflow (on: workflow_call) and call it from
velox_backend_x86.yml so the Delta tests reuse the native lib + arrow jars that
workflow already builds, instead of duplicating the build-native-lib-centos-7
job. GitHub artifacts cannot be shared across workflows, so the only way to
reuse the artifact is to run the Delta jobs in the same workflow run.

delta_spark_ut.yml keeps a workflow_dispatch trigger for standalone manual runs
(its build-native-lib-centos-7 job is gated to that case and skipped when
called); the pull_request trigger is removed so the suite no longer double-runs.
velox_backend_x86.yml gains an arrow-jars upload on its native build and a
delta-spark-ut job that calls the reusable workflow. That job runs on every
velox trigger like the other spark-test jobs, since core/velox/substrait/cpp
changes can affect Delta query offload.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

[CI] Harden Delta UT setup clone and enforce-mode baseline check

Address PR review feedback:

- setup-delta.sh: replace the shallow-clone + full-clone fallback (which ran a
  destructive `rm -rf "$DELTA_DIR"`) with a single `git init` + shallow
  `fetch --depth 1 origin "$DELTA_REF"` + `checkout FETCH_HEAD`. This resolves a
  tag, branch, or commit SHA uniformly (`git clone --branch` rejects SHAs),
  drops the dead fallback branch, and removes the unguarded recursive delete.

- compare-test-results.py: in enforce mode, a missing/typoed --known-failures
  path made load_entries() return an empty set, silently degrading to seed mode
  and passing the gate without enforcing regressions. Treat a missing baseline
  file as a configuration error (exit 2); an existing-but-empty file is still
  allowed and legitimately seeds.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

[CI] Harden Delta UT gate and setup against silent failures

Address PR review feedback with four robustness fixes:

- compare-test-results.py (enforce/seed): raise NoReportsError and exit 2 when no
  JUnit <testsuite> elements are parsed, instead of warning and returning empty
  sets. Otherwise a misconfiguration (wrong reports dir, broken reporter, suites
  crashing before writing XML) yields zero failures -> zero regressions -> a
  silent green gate.
- compare-test-results.py (aggregate): exit 2 before writing baseline-out when no
  per-shard failures-*.txt / ran-*.txt inputs are found. The gate-list download is
  continue-on-error and aggregate runs with if: always(), so missing artifacts
  would otherwise produce an empty baseline that could be committed, wiping
  known-failures.txt.
- setup-delta.sh: pass the Delta ref after `--` in git fetch so a ref starting
  with `-` can't be misread as a git option (the script is workflow_dispatch-
  runnable with a user-supplied ref).
- velox_backend_x86.yml: drop secrets: inherit from the reusable Delta UT call.
  delta_spark_ut.yml references no secrets, so inheriting them needlessly forwards
  all caller secrets to a workflow that clones and runs external code.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

[CI] Quarantine flaky Delta DV-merge tests in the UT gate

Some Delta-on-Gluten MERGE tests that write deletion vectors fail
non-deterministically: the same bundle passes them on one CI run and
fails them on the next (native RoaringBitmapArray addSafe aborting on an
invalid Long.MAX_VALUE row index). Such tests cannot live in
known-failures.txt -- baselining them reds the gate on every run where
they pass, and leaving them out reds it on every run where they fail.

Add a flaky-tests.txt quarantine list read by the gate. A quarantined
test is neutral: it never counts as a regression when it fails nor as
now-passing when it passes, and is excluded from the regenerated
baseline (aggregate mode). The suite portion of each entry is an fnmatch
glob so one line covers a root-cause family across generated suite
variants (e.g. *DVs*Suite); the test name is matched exactly.

Seed the list with the DV-merge family behind the native row-index bug.
This is an interim measure -- entries should be removed once that bug is
fixed in the native backend so the tests are enforced again.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

[CI] Address Delta UT review: single-source spark_version, corrupt-report fail-fast, baseline count

Fixes three open review comments on the Delta Spark UT pipeline:

1. spark_version was redundant with the hard-coded GLUTEN_BUNDLE_SPARK_VERSION /
   GLUTEN_SPARK_PROFILE, and a workflow_dispatch run could set them out of sync so
   the Delta tests ran against a mismatched Gluten bundle. Make spark_version the
   single source of truth: the Gluten bundle profile (-Pspark-<v>), the bundle jar
   name and Delta's -DsparkVersion are all derived from it, so a mismatch is now
   impossible (no guard needed). Scala 2.13 / JDK 17 stay pinned.

2. Corrupt/truncated JUnit reports (compare-test-results.py). parse_reports
   previously warned and skipped any XML that failed to parse, so a report
   truncated by a killed/OOM'd fork could silently drop a suite's failures and
   let the gate go green on partial data. A TEST-*.xml that fails to parse now
   raises CorruptReportError (exit 2); other XML matched by the broad target/**
   glob is still skipped.

3. Hard-coded failure count in the baseline header (known-failures.txt). The
   header stated a fixed total that drifts every time the baseline is refreshed.
   Drop the number and point at the entry count / delta-spark-aggregate summary
   as the authoritative source instead.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

[CI] Fail Delta UT aggregate when a shard's gate lists are missing

The aggregate job downloads per-shard gate lists with continue-on-error and the
aggregator only hard-failed when *no* inputs were found. If a shard died before
writing its gate lists (OOM / watchdog) or its artifact failed to download, the
job still regenerated a baseline that silently omitted that shard's failures,
shrinking known-failures.txt and reddening the next run.

Add --expected-shards to compare-test-results.py: in aggregate mode it counts
shards that produced a complete failures-*/ran-* pair and refuses to aggregate
(exit 2, before writing --baseline-out) when fewer than expected are present. A
shard counts only when BOTH files exist, so a partial download is caught too.
The workflow passes --expected-shards ${{ env.DELTA_NUM_SHARDS }}; omitting the
flag (default 0) disables the check, preserving prior behavior.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

[CI] Slim Delta UT pipeline: drop redundant Arrow jars, extract shard body

Two readability/cleanup changes to the Delta Spark UT pipeline:

1. Drop redundant Arrow jar handling. Since apache#12244
   (2026-06-09) Gluten depends on vanilla Apache Arrow instead of the
   custom 15.0.0-gluten rename, so the Arrow jars resolve from Maven
   Central. Remove the arrow-jar staging in both native builds, the
   velox-arrow-jars / delta-spark-ut-arrow-jars uploads, the caller's
   arrow_jars_artifact input, and the bundle job's "Download Arrow jars"
   step. The bundle build's Maven cache still holds Arrow across runs and
   resolves it from Central on a miss.

2. Extract the shard test body into run-delta-tests.sh. The "Run Delta
   spark module tests" step embedded ~190 lines of shell (hang watchdog,
   sbt invocation with JVM/heap tuning, cgroup memory forensics, report
   check, known-failures gate). Move that whole body into
   util/delta-spark-ut/run-delta-tests.sh so the workflow step is just an
   `env:` block plus a one-line script call, shrinking delta_spark_ut.yml
   by ~180 lines. The move is faithful: the script body is the previous
   inline block with only the GitHub `${{ }}` expressions replaced by env
   vars (matrix.shard -> SHARD_ID which is already job-level env;
   spark_version/update_baseline/fail_on_fixed added to the step env).
   Verified by a textual diff against the old body and by end-to-end runs
   with a stubbed sbt (baseline-only -> green, regression -> red, seed ->
   green).

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

[CI] Share Delta test JVM flags via java-test-args.sh for local runs

The Gluten/JDK17 test JVM flags (--add-opens + the Netty reflection
property, mirroring root pom.xml's extraJavaTestArgs) lived only in the
workflow step's env: block, so a developer running the Delta suite
locally had no shared source for them.

Move them into util/delta-spark-ut/java-test-args.sh, which exports
JAVA_TOOL_OPTIONS. run-delta-tests.sh sources it (via a BASH_SOURCE-
relative path so it also works for local runs), and the workflow env
block no longer sets JAVA_TOOL_OPTIONS. A developer can source the same
file before `sbt spark/test` to get identical flags; README documents it.

Faithful: the sourced value is byte-identical to the previous env-block
value (verified), and an end-to-end run with a stubbed sbt confirms the
forked test JVM still receives all 18 flags with JAVA_TOOL_OPTIONS unset
in the environment.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

[CI] Quarantine flaky Delta failures by error signature, not test name

The native Delta DV bitmap row-index bug (RoaringBitmapArray aborting on a
Long.MAX_VALUE row index during a MERGE that writes deletion vectors) is
intermittent and lands on a DIFFERENT *DVs*Suite MERGE test each run, so
listing every test in flaky-tests.txt was whack-a-mole (6 entries and
counting).

Add error-signature quarantine: the gate now records each failed test's
JUnit <failure>/<error> text, and flaky-error-patterns.txt holds regexes
matched against it. A failure matching a pattern is treated as flaky
regardless of which test it hit -- neither counted as a regression nor
written to the shard's failures list (so it can't leak into the
regenerated baseline). Seeded with the RoaringBitmapArray signature.

This is more precise than a name glob: a different real failure in the
same DV suite is still caught, because only failures carrying the
signature are ignored. The six *DVs*Suite name entries are removed from
flaky-tests.txt (superseded); the name mechanism stays for flakes without
a distinctive error signature.

Verified against a real failing shard report: the DV failure is a
regression without the patterns file and quarantined (gate green) with it.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

[CI] Quarantine the negative-index DV-bitmap flaky variant

Run 29042495519 regressed on another DV-merge test whose failure was NOT
the RoaringBitmapArray "exceeds max representable value" signature but a
second, distinct native error from the same DV bitmap aggregation --

  Reason: Delta bitmap row index cannot be negative: -6254810385378525259
  Expression: value >= 0
  Function: addRowIndex  File: .../delta/DeltaBitmapAggregator.cc:44

vs the previously-seen too-large variant (value <= kMaxRepresentableValue,
RoaringBitmapArray.cpp addSafe). Same root cause (garbage row index into
the DV bitmap aggregator), different bounds check and different message.

Add a second explicit pattern for it rather than broadening the existing
one -- the two are distinct native error strings, so keeping each pattern
tightly bound to its message avoids masking an unrelated failure that
merely mentions a bitmap row index. Verified against the real failing
report: the negative-index failure is a regression without the pattern and
quarantined with it; the suite's five other (non-bitmap) baseline failures
and benign controls are not matched.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

[CI] Remove now-passing Variant/UDT tests from Delta baseline

The Velox Variant/UDT type-offload gap was fixed upstream, so 74 Delta UT
tests (65 Variant + 9 user-defined-type) that were expected failures now pass
and were tripping the fail-on-fixed gate as "now-passing".

Remove them from known-failures.txt (810 -> 736). The 6 remaining
`variant auto compact` entries are a different root cause (auto-compact
metrics) and correctly stay in the baseline.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Copilot-Session: 66a9e40f-3ac8-45be-8fee-a606a22fa098

[CI] Remove now-passing DV-tombstone test from Delta baseline

DeltaFastDropFeatureSuite "We do not create redundant DV tombstones after
cloning isShallowClone: true" now passes consistently (green on 2026-07-13 and
2026-07-14 runs); the earlier FileNotFoundException on a deletion-vector .bin
during the DROP FEATURE parallel OPTIMIZE no longer reproduces. Keeping it in
the baseline trips the fail-on-fixed gate on every run, so remove it (736 ->
735).

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Copilot-Session: 66a9e40f-3ac8-45be-8fee-a606a22fa098

[CI] Fail Delta shard on watchdog kill; kill only forks; normalize gate keys

Addresses three review comments on the Delta Spark UT pipeline:

* Hang watchdog: when it kills a wedged test fork, the running suite plus
  every suite queued behind it in that JVM never run and never write a
  report; since we ignore sbt's exit code, the gate only judged the suites
  that reported and the shard could go green. The watchdog now touches a
  marker on kill and the shard fails afterwards if it exists.

* Hang watchdog: only KILL matched sbt.ForkMain fork(s), never the sbt
  launcher. Before any fork exists (dependency resolution / cold-cache
  compile) sbt can be silent for >15 min; killing the launcher then wasted
  the slot on a confusing "compile/launch failure". All JVMs are still
  dumped for diagnostics. The per-episode dump/kill budget resets when
  output resumes so a transient pre-fork stall can't starve a later fork
  hang.

* Gate: normalize (suite, test) keys parsed from JUnit XML the same way
  baseline/flaky entries are (write_entries collapses CR/LF; parse_entry
  strips the line), so a test name with a trailing newline or surrounding
  whitespace is suppressible by a line pasted from the gate's REGRESSION
  output.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Copilot-Session: 66a9e40f-3ac8-45be-8fee-a606a22fa098

[CI] Gate Delta UT per-PR to Delta paths + label; add nightly full run

Reduce GitHub Actions usage for the Delta Spark UT suite (per review feedback
on the pipeline PR):

* Per PR (velox_backend_x86.yml): a new `delta-changes` job runs the suite only
  when the PR touches high-signal Delta paths -- the Delta integration code
  (backends-velox/src-delta*), the gluten-delta module, or this pipeline's own
  files -- or carries the `run-delta-ci` opt-in label. Changes to general
  Velox/core/native code (touched on most PRs) skip it; this drops the per-PR
  trigger rate from ~60% to ~17% of recent commits.

* Nightly (delta_spark_ut.yml): a `schedule` (05:00 UTC) runs the full suite
  against the latest default branch, so regressions from the skipped-per-PR
  paths are still caught daily. It builds its own native lib (no caller) and
  uses fail_on_fixed=true, so baseline drift surfaces as a red nightly -- the
  signal to refresh known-failures.txt.

* Docs: README "When it runs" section documents the per-PR path gate, the
  opt-in label, and the nightly run.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Copilot-Session: 66a9e40f-3ac8-45be-8fee-a606a22fa098

[CI] Fix stale Delta-job comment in velox_backend_x86.yml

The comment above the Delta Spark UT job still said it "runs on every trigger
like the other spark-test jobs", which contradicted the per-PR gating added by
the delta-changes job. Remove the stale block; keep the accurate gating comment
and add a concise native-lib-reuse note on the delta-spark-ut job.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Copilot-Session: 66a9e40f-3ac8-45be-8fee-a606a22fa098

[CI] Baseline new ImplicitStreamingMergeCasting BIGINT->DECIMAL overflow failure

After the latest rebase, ImplicitStreamingMergeCastingSuite "Streaming MERGE
overflow sourceType: BIGINT, targetType: DECIMAL(7,2) followAnsiEnabled: false,
ansiEnabled: true, storeAssignmentPolicy: LEGACY" fails deterministically:
Velox raises `VeloxUserError INVALID_ARGUMENT: Cannot cast BIGINT
'9223372036854775807' to DECIMAL(7,2)` (rescaleInt, DecimalUtil.h) where vanilla
Spark handles the overflow per the ANSI/LEGACY policy. Its followAnsiEnabled:true
siblings are already baselined; add this variant so the gate stays green
(735 -> 736 entries, still sorted).

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Copilot-Session: 66a9e40f-3ac8-45be-8fee-a606a22fa098

[CI] Skip already-applied Delta cherry-picks in setup-delta.sh

setup-delta.sh runs under `set -euo pipefail`, so cherry-picking a Delta
FileSourceScanLike fix that the pinned DELTA_REF already contains exits
non-zero (empty/conflicting patch) and aborts the whole setup. This is a
latent break the moment DELTA_REF is bumped past delta-io/delta#7104/apache#7105.

Make cherry_pick_delta_fix attempt the cherry-pick and, on failure, recover
only the paths that fix touches (git diff-tree -> per-file reset + checkout,
then cherry-pick --quit) and continue -- leaving the DeltaSQLCommandTest
patch and bundle jar intact. It is self-correcting: a genuinely missing fix
resurfaces as gate regressions instead of a hard abort.

Ancestry can't distinguish "already contained" from a real conflict here
because the Delta clone is shallow (depth 1) and merge-base --is-ancestor
can't see past the graft; a reverse-apply-check is fragile when the newer
ref carries the fix plus adjacent edits. Recover-on-failure avoids both.

Generated-by: GitHub Copilot CLI
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Copilot-Session: 66a9e40f-3ac8-45be-8fee-a606a22fa098

[CI] Track skipped Delta tests separately and make the clone step idempotent

Two fixes from review feedback on the Delta gate.

1. compare-test-results.py treated a skipped test as "not seen this run", so a
   baseline entry that merely got skipped was reported under "Stale baseline
   entries (suite/test gone)" and silently dropped from the regenerated
   baseline -- only to return as a regression the next time it executed.

   Each shard now writes a --skipped-out list alongside --ran-out, and the
   aggregate job reports those entries as "Skipped this run" instead of stale
   and carries them over into the regenerated baseline. Genuinely removed tests
   appear in neither list and are still reported as stale.

   Skips are deliberately NOT folded into --ran-out: "now-passing" is derived
   from `ran - failed`, so counting a skip as a run would report it as fixed
   and, under fail_on_fixed, demand its removal from the baseline. Missing
   skipped-*.txt (older artifacts) degrades to the previous behaviour, and the
   shard-completeness guard still keys off failures-*/ran-* only.

2. setup-delta.sh could not be re-run over an existing DELTA_DIR: `git remote
   add origin` exits 3 when origin already exists, which aborts the script
   under `set -euo pipefail` and forces manual cleanup. Drop the remote first
   and force the checkout so a partial previous run is recovered rather than
   fatal. No `rm -rf` is reintroduced; the bundle jar and source patches are
   applied after this block, so nothing worth keeping is discarded.

Generated-by: GitHub Copilot CLI
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Copilot-Session: 66a9e40f-3ac8-45be-8fee-a606a22fa098

[CI] Report invalid flaky-error regexes clearly and tighten the DV pattern

Follow-ups from review feedback on the Delta gate.

- load_patterns() compiled each line of flaky-error-patterns.txt with no error
  handling, so a typo in that hand-edited file surfaced as a bare re.error
  traceback naming neither the file nor the offending line. Raise
  BadPatternError with path, line number and pattern, and exit 2 from the gate:
  "<path>:4: invalid regex '[unclosed': unterminated character set at position 0".

- The negative-row-index signature used `-?\d+`, which would also match a
  positive index. The native check is `value >= 0`, so the reported index is
  always negative; match the sign explicitly, per this file's own rule of
  binding patterns to the exact error string so a quarantine can't swallow an
  unrelated failure.

- Note in the DeletionVectorsSuite block that the sed depends on the clone
  step's `checkout -f`. The sed appends after the test-declaration line, so
  without that per-run reset a re-run injects duplicate `fail` lines and trips
  the INJECTED != 2 check (verified: 2, 4, 6 without `-f`; 2, 2, 2 with it).

Generated-by: GitHub Copilot CLI
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Copilot-Session: 66a9e40f-3ac8-45be-8fee-a606a22fa098

[CI] Make the Delta path gate fail open when git diff fails

The delta-changes job intends to run the Delta suite whenever it cannot
determine what the PR touched -- "never silently skip coverage". But the
detection itself was fail-closed:

  if git diff --name-only "$BASE" "$HEAD_SHA" | grep -Eq '<delta paths>'; then

When git diff fails (missing objects after a force-push race, an unfetched
fork head, or the `git merge-base ... || echo "$BASE_SHA"` fallback above
handing it an unusable sha), grep gets empty input and exits 1 -- exactly as
it does for "no Delta paths changed" -- so the else branch set run_delta=false
and the suite was skipped. `set -euo pipefail` does not help here: a command
used as an `if` condition is allowed to fail.

Capture the diff first and treat a git failure as an explicit fail-open, then
match against the captured output.

Verified over 8 scenarios: bad base+head, valid base + bad head and empty shas
now all yield run_delta=true, while gluten-delta/, backends-velox/src-delta40,
docs-only, cpp/velox-only and no-change diffs are unchanged. The previous
logic fails the first three.

Generated-by: GitHub Copilot CLI
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Copilot-Session: 66a9e40f-3ac8-45be-8fee-a606a22fa098

[CI] Baseline new ImplicitMergeCasting BIGINT->DECIMAL overflow failure

Shard 2 of run 30324192115 reports one regression:

  ImplicitMergeCastingSuite#MERGE overflow in WHEN MATCHED THEN UPDATE SET
  t.value = s.value sourceType: BIGINT, targetType: DECIMAL(7,2)
  followAnsiEnabled: false, ansiEnabled: true, storeAssignmentPolicy: LEGACY

Under storeAssignmentPolicy LEGACY the overflowing value is expected to be
stored without error, but the offloaded Velox cast raises instead:

  VeloxUserError INVALID_ARGUMENT
  Reason: Cannot cast BIGINT '9223372036854775807' to DECIMAL(7, 2)
  Function: rescaleInt  File: velox/type/DecimalUtil.h:218

This is the non-streaming sibling of the ImplicitStreamingMergeCastingSuite
case for the same BIGINT -> DECIMAL(7,2) / LEGACY combination that was
baselined earlier; both surfaced after the rebase. It went unnoticed one run
longer because shard 2 of run 30186097028 died on a Docker pull before the
gate ran, so the test never executed there.

Baselined rather than quarantined in flaky-tests.txt: it failed in both runs
where it actually executed, with an identical error, and the failure is a pure
expression-evaluation overflow with no dependence on scheduling or runtime
plan -- unlike the DV bitmap row-index bug that flaky-error-patterns.txt
covers. Keeping it in the baseline means the gate will tell us to remove it
once the LEGACY cast semantics are fixed, which a flaky entry would not.

736 -> 737 entries.

Generated-by: GitHub Copilot CLI
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Copilot-Session: 66a9e40f-3ac8-45be-8fee-a606a22fa098
felipepessoto added a commit to felipepessoto/gluten that referenced this pull request Jul 30, 2026
… body

Two readability/cleanup changes to the Delta Spark UT pipeline:

1. Drop redundant Arrow jar handling. Since apache#12244
   (2026-06-09) Gluten depends on vanilla Apache Arrow instead of the
   custom 15.0.0-gluten rename, so the Arrow jars resolve from Maven
   Central. Remove the arrow-jar staging in both native builds, the
   velox-arrow-jars / delta-spark-ut-arrow-jars uploads, the caller's
   arrow_jars_artifact input, and the bundle job's "Download Arrow jars"
   step. The bundle build's Maven cache still holds Arrow across runs and
   resolves it from Central on a miss.

2. Extract the shard test body into run-delta-tests.sh. The "Run Delta
   spark module tests" step embedded ~190 lines of shell (hang watchdog,
   sbt invocation with JVM/heap tuning, cgroup memory forensics, report
   check, known-failures gate). Move that whole body into
   util/delta-spark-ut/run-delta-tests.sh so the workflow step is just an
   `env:` block plus a one-line script call, shrinking delta_spark_ut.yml
   by ~180 lines. The move is faithful: the script body is the previous
   inline block with only the GitHub `${{ }}` expressions replaced by env
   vars (matrix.shard -> SHARD_ID which is already job-level env;
   spark_version/update_baseline/fail_on_fixed added to the step env).
   Verified by a textual diff against the old body and by end-to-end runs
   with a stubbed sbt (baseline-only -> green, regression -> red, seed ->
   green).

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
zhouyuan pushed a commit that referenced this pull request Aug 10, 2026
…s baseline (#12388)

* [GLUTEN][CI] Add Delta Spark UT pipeline gated against a known-failures baseline

Run delta-io/delta's `spark` ScalaTest suite against a Gluten Velox bundle in CI
and gate the results against a committed baseline so the many expected Delta-on-
Gluten failures stay manageable and can be fixed incrementally without letting
currently-passing tests silently regress.

What it adds (.github/workflows/util/delta-spark-ut/):
- delta_spark_ut.yml: builds the native lib + Gluten bundle, then runs the Delta
  spark suite sharded by suite into 4 shards x 4 forked test JVMs (~16-way), and
  gates each shard against the baseline.
- compare-test-results.py: the gate. Per shard, regressions (failed not in the
  baseline) fail the build; newly-passing baselined tests are flagged so the
  baseline can be tightened. Also supports seed/aggregate modes.
- known-failures.txt: the committed baseline of expected failures.
- setup-delta.sh: clones Delta, injects the Gluten bundle, patches
  DeltaSQLCommandTest, and force-fails the two DeletionVectorsSuite 2B-row tests
  whose native row-index materialization OOM-kills the runner and hangs the shard.
- README.md: how the pipeline, gating and baseline-refresh work.

The workflow also carries a hang watchdog that thread-dumps and kills a wedged
fork, and tunes the per-fork heap (2G) and off-heap (2G) to fit the ~16G runner.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* Change order of steps

Co-authored-by: Copilot Autofix powered by AI <175728472+Copilot@users.noreply.github.com>

* [CI] Cherry-pick Delta FileSourceScanLike test fixes and refresh baseline

Delta's data-skipping, limit-push-down, column-pruning and scan-metric tests
collect file-source scans by matching the concrete `FileSourceScanExec` case
class. Under the Gluten Velox bundle the scan is offloaded to
DeltaScanTransformer, a sibling that implements the same `FileSourceScanLike`
interface but is not FileSourceScanExec, so the match misses and the scan
looks absent. This surfaced as `scala.MatchError: List()` (~56
DataSkipping*/DeltaLimitPushDown* tests), empty generated-column partition
filters (~45 OptimizeGeneratedColumnSuite tests) and broken column-pruning /
scan-metric checks across the Delete, Update, Merge, DeletionVectors and
RowId suites and the TestsStatistics helper.

Gluten copies `partitionFilters` and the other accessors these tests read
verbatim onto the offloaded scan, so results are identical to vanilla -- only
the test's `case` match breaks. Fix it by cherry-picking the two merged
upstream Delta commits that widen these matches to the shared
`FileSourceScanLike` interface (behavior-preserving for vanilla, which also
implements it):

  * delta-io/delta#7104 -- ScanReportHelper.collectScans
  * delta-io/delta#7105 -- the remaining 9 test sources, its follow-up

Both are merged on Delta master but land after the ref this workflow builds
against (v4.2.0), so setup-delta.sh cherry-picks them onto the shallow
checkout. Each fetches the fix commit at depth 2 (commit + parent) so
cherry-pick can compute the parent->fix diff, and uses `cherry-pick -n` so no
committer identity is required. Once the pinned DELTA_REF advances to include
a commit its cherry-pick becomes a clean no-op and that block can be removed.

The cherry-picks run before the DeletionVectorsSuite 2B-row force-fail step:
that step sed-injects fail() into DeletionVectorsSuite.scala, which
delta-io/delta#7105 also edits, and git cherry-pick refuses to apply onto a
working tree with uncommitted changes to a file it touches (exit 128).

Refresh known-failures.txt from run 28299900971 (the delta-spark-aggregate job
output), which ran all 19073 tests across 16 shards: removes 187 now-passing
tests with 0 regressions, 963 -> 776. ~147 come from the fixes above
(DataSkipping*, DeltaLimitPushDown*, OptimizeGeneratedColumnSuite, MergeInto*,
RowIdSuite); the remaining ~40 are other suites that now pass (e.g.
HiveConvertToDeltaSuite, BitmapAggregatorE2ESuite). Verified against the
per-shard ran/failed lists: every baseline entry was observed this run (0
stale), so nothing was dropped due to a crashed or incomplete shard.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* [CI] Reuse velox_backend_x86's native build for Delta Spark UT

Make delta_spark_ut.yml a reusable workflow (on: workflow_call) and call it from
velox_backend_x86.yml so the Delta tests reuse the native lib + arrow jars that
workflow already builds, instead of duplicating the build-native-lib-centos-7
job. GitHub artifacts cannot be shared across workflows, so the only way to
reuse the artifact is to run the Delta jobs in the same workflow run.

delta_spark_ut.yml keeps a workflow_dispatch trigger for standalone manual runs
(its build-native-lib-centos-7 job is gated to that case and skipped when
called); the pull_request trigger is removed so the suite no longer double-runs.
velox_backend_x86.yml gains an arrow-jars upload on its native build and a
delta-spark-ut job that calls the reusable workflow. That job runs on every
velox trigger like the other spark-test jobs, since core/velox/substrait/cpp
changes can affect Delta query offload.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* [CI] Harden Delta UT setup clone and enforce-mode baseline check

Address PR review feedback:

- setup-delta.sh: replace the shallow-clone + full-clone fallback (which ran a
  destructive `rm -rf "$DELTA_DIR"`) with a single `git init` + shallow
  `fetch --depth 1 origin "$DELTA_REF"` + `checkout FETCH_HEAD`. This resolves a
  tag, branch, or commit SHA uniformly (`git clone --branch` rejects SHAs),
  drops the dead fallback branch, and removes the unguarded recursive delete.

- compare-test-results.py: in enforce mode, a missing/typoed --known-failures
  path made load_entries() return an empty set, silently degrading to seed mode
  and passing the gate without enforcing regressions. Treat a missing baseline
  file as a configuration error (exit 2); an existing-but-empty file is still
  allowed and legitimately seeds.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* [CI] Harden Delta UT gate and setup against silent failures

Address PR review feedback with four robustness fixes:

- compare-test-results.py (enforce/seed): raise NoReportsError and exit 2 when no
  JUnit <testsuite> elements are parsed, instead of warning and returning empty
  sets. Otherwise a misconfiguration (wrong reports dir, broken reporter, suites
  crashing before writing XML) yields zero failures -> zero regressions -> a
  silent green gate.
- compare-test-results.py (aggregate): exit 2 before writing baseline-out when no
  per-shard failures-*.txt / ran-*.txt inputs are found. The gate-list download is
  continue-on-error and aggregate runs with if: always(), so missing artifacts
  would otherwise produce an empty baseline that could be committed, wiping
  known-failures.txt.
- setup-delta.sh: pass the Delta ref after `--` in git fetch so a ref starting
  with `-` can't be misread as a git option (the script is workflow_dispatch-
  runnable with a user-supplied ref).
- velox_backend_x86.yml: drop secrets: inherit from the reusable Delta UT call.
  delta_spark_ut.yml references no secrets, so inheriting them needlessly forwards
  all caller secrets to a workflow that clones and runs external code.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* [CI] Quarantine flaky Delta DV-merge tests in the UT gate

Some Delta-on-Gluten MERGE tests that write deletion vectors fail
non-deterministically: the same bundle passes them on one CI run and
fails them on the next (native RoaringBitmapArray addSafe aborting on an
invalid Long.MAX_VALUE row index). Such tests cannot live in
known-failures.txt -- baselining them reds the gate on every run where
they pass, and leaving them out reds it on every run where they fail.

Add a flaky-tests.txt quarantine list read by the gate. A quarantined
test is neutral: it never counts as a regression when it fails nor as
now-passing when it passes, and is excluded from the regenerated
baseline (aggregate mode). The suite portion of each entry is an fnmatch
glob so one line covers a root-cause family across generated suite
variants (e.g. *DVs*Suite); the test name is matched exactly.

Seed the list with the DV-merge family behind the native row-index bug.
This is an interim measure -- entries should be removed once that bug is
fixed in the native backend so the tests are enforced again.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* [CI] Address Delta UT review: single-source spark_version, corrupt-report fail-fast, baseline count

Fixes three open review comments on the Delta Spark UT pipeline:

1. spark_version was redundant with the hard-coded GLUTEN_BUNDLE_SPARK_VERSION /
   GLUTEN_SPARK_PROFILE, and a workflow_dispatch run could set them out of sync so
   the Delta tests ran against a mismatched Gluten bundle. Make spark_version the
   single source of truth: the Gluten bundle profile (-Pspark-<v>), the bundle jar
   name and Delta's -DsparkVersion are all derived from it, so a mismatch is now
   impossible (no guard needed). Scala 2.13 / JDK 17 stay pinned.

2. Corrupt/truncated JUnit reports (compare-test-results.py). parse_reports
   previously warned and skipped any XML that failed to parse, so a report
   truncated by a killed/OOM'd fork could silently drop a suite's failures and
   let the gate go green on partial data. A TEST-*.xml that fails to parse now
   raises CorruptReportError (exit 2); other XML matched by the broad target/**
   glob is still skipped.

3. Hard-coded failure count in the baseline header (known-failures.txt). The
   header stated a fixed total that drifts every time the baseline is refreshed.
   Drop the number and point at the entry count / delta-spark-aggregate summary
   as the authoritative source instead.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* [CI] Fail Delta UT aggregate when a shard's gate lists are missing

The aggregate job downloads per-shard gate lists with continue-on-error and the
aggregator only hard-failed when *no* inputs were found. If a shard died before
writing its gate lists (OOM / watchdog) or its artifact failed to download, the
job still regenerated a baseline that silently omitted that shard's failures,
shrinking known-failures.txt and reddening the next run.

Add --expected-shards to compare-test-results.py: in aggregate mode it counts
shards that produced a complete failures-*/ran-* pair and refuses to aggregate
(exit 2, before writing --baseline-out) when fewer than expected are present. A
shard counts only when BOTH files exist, so a partial download is caught too.
The workflow passes --expected-shards ${{ env.DELTA_NUM_SHARDS }}; omitting the
flag (default 0) disables the check, preserving prior behavior.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* [CI] Slim Delta UT pipeline: drop redundant Arrow jars, extract shard body

Two readability/cleanup changes to the Delta Spark UT pipeline:

1. Drop redundant Arrow jar handling. Since #12244
   (2026-06-09) Gluten depends on vanilla Apache Arrow instead of the
   custom 15.0.0-gluten rename, so the Arrow jars resolve from Maven
   Central. Remove the arrow-jar staging in both native builds, the
   velox-arrow-jars / delta-spark-ut-arrow-jars uploads, the caller's
   arrow_jars_artifact input, and the bundle job's "Download Arrow jars"
   step. The bundle build's Maven cache still holds Arrow across runs and
   resolves it from Central on a miss.

2. Extract the shard test body into run-delta-tests.sh. The "Run Delta
   spark module tests" step embedded ~190 lines of shell (hang watchdog,
   sbt invocation with JVM/heap tuning, cgroup memory forensics, report
   check, known-failures gate). Move that whole body into
   util/delta-spark-ut/run-delta-tests.sh so the workflow step is just an
   `env:` block plus a one-line script call, shrinking delta_spark_ut.yml
   by ~180 lines. The move is faithful: the script body is the previous
   inline block with only the GitHub `${{ }}` expressions replaced by env
   vars (matrix.shard -> SHARD_ID which is already job-level env;
   spark_version/update_baseline/fail_on_fixed added to the step env).
   Verified by a textual diff against the old body and by end-to-end runs
   with a stubbed sbt (baseline-only -> green, regression -> red, seed ->
   green).

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* [CI] Share Delta test JVM flags via java-test-args.sh for local runs

The Gluten/JDK17 test JVM flags (--add-opens + the Netty reflection
property, mirroring root pom.xml's extraJavaTestArgs) lived only in the
workflow step's env: block, so a developer running the Delta suite
locally had no shared source for them.

Move them into util/delta-spark-ut/java-test-args.sh, which exports
JAVA_TOOL_OPTIONS. run-delta-tests.sh sources it (via a BASH_SOURCE-
relative path so it also works for local runs), and the workflow env
block no longer sets JAVA_TOOL_OPTIONS. A developer can source the same
file before `sbt spark/test` to get identical flags; README documents it.

Faithful: the sourced value is byte-identical to the previous env-block
value (verified), and an end-to-end run with a stubbed sbt confirms the
forked test JVM still receives all 18 flags with JAVA_TOOL_OPTIONS unset
in the environment.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* [CI] Quarantine flaky Delta failures by error signature, not test name

The native Delta DV bitmap row-index bug (RoaringBitmapArray aborting on a
Long.MAX_VALUE row index during a MERGE that writes deletion vectors) is
intermittent and lands on a DIFFERENT *DVs*Suite MERGE test each run, so
listing every test in flaky-tests.txt was whack-a-mole (6 entries and
counting).

Add error-signature quarantine: the gate now records each failed test's
JUnit <failure>/<error> text, and flaky-error-patterns.txt holds regexes
matched against it. A failure matching a pattern is treated as flaky
regardless of which test it hit -- neither counted as a regression nor
written to the shard's failures list (so it can't leak into the
regenerated baseline). Seeded with the RoaringBitmapArray signature.

This is more precise than a name glob: a different real failure in the
same DV suite is still caught, because only failures carrying the
signature are ignored. The six *DVs*Suite name entries are removed from
flaky-tests.txt (superseded); the name mechanism stays for flakes without
a distinctive error signature.

Verified against a real failing shard report: the DV failure is a
regression without the patterns file and quarantined (gate green) with it.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* [CI] Quarantine the negative-index DV-bitmap flaky variant

Run 29042495519 regressed on another DV-merge test whose failure was NOT
the RoaringBitmapArray "exceeds max representable value" signature but a
second, distinct native error from the same DV bitmap aggregation --

  Reason: Delta bitmap row index cannot be negative: -6254810385378525259
  Expression: value >= 0
  Function: addRowIndex  File: .../delta/DeltaBitmapAggregator.cc:44

vs the previously-seen too-large variant (value <= kMaxRepresentableValue,
RoaringBitmapArray.cpp addSafe). Same root cause (garbage row index into
the DV bitmap aggregator), different bounds check and different message.

Add a second explicit pattern for it rather than broadening the existing
one -- the two are distinct native error strings, so keeping each pattern
tightly bound to its message avoids masking an unrelated failure that
merely mentions a bitmap row index. Verified against the real failing
report: the negative-index failure is a regression without the pattern and
quarantined with it; the suite's five other (non-bitmap) baseline failures
and benign controls are not matched.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* [CI] Remove now-passing Variant/UDT tests from Delta baseline

The Velox Variant/UDT type-offload gap was fixed upstream, so 74 Delta UT
tests (65 Variant + 9 user-defined-type) that were expected failures now pass
and were tripping the fail-on-fixed gate as "now-passing".

Remove them from known-failures.txt (810 -> 736). The 6 remaining
`variant auto compact` entries are a different root cause (auto-compact
metrics) and correctly stay in the baseline.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Copilot-Session: 66a9e40f-3ac8-45be-8fee-a606a22fa098

* [CI] Remove now-passing DV-tombstone test from Delta baseline

DeltaFastDropFeatureSuite "We do not create redundant DV tombstones after
cloning isShallowClone: true" now passes consistently (green on 2026-07-13 and
2026-07-14 runs); the earlier FileNotFoundException on a deletion-vector .bin
during the DROP FEATURE parallel OPTIMIZE no longer reproduces. Keeping it in
the baseline trips the fail-on-fixed gate on every run, so remove it (736 ->
735).

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Copilot-Session: 66a9e40f-3ac8-45be-8fee-a606a22fa098

* [CI] Fail Delta shard on watchdog kill; kill only forks; normalize gate keys

Addresses three review comments on the Delta Spark UT pipeline:

* Hang watchdog: when it kills a wedged test fork, the running suite plus
  every suite queued behind it in that JVM never run and never write a
  report; since we ignore sbt's exit code, the gate only judged the suites
  that reported and the shard could go green. The watchdog now touches a
  marker on kill and the shard fails afterwards if it exists.

* Hang watchdog: only KILL matched sbt.ForkMain fork(s), never the sbt
  launcher. Before any fork exists (dependency resolution / cold-cache
  compile) sbt can be silent for >15 min; killing the launcher then wasted
  the slot on a confusing "compile/launch failure". All JVMs are still
  dumped for diagnostics. The per-episode dump/kill budget resets when
  output resumes so a transient pre-fork stall can't starve a later fork
  hang.

* Gate: normalize (suite, test) keys parsed from JUnit XML the same way
  baseline/flaky entries are (write_entries collapses CR/LF; parse_entry
  strips the line), so a test name with a trailing newline or surrounding
  whitespace is suppressible by a line pasted from the gate's REGRESSION
  output.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Copilot-Session: 66a9e40f-3ac8-45be-8fee-a606a22fa098

* [CI] Gate Delta UT per-PR to Delta paths + label; add nightly full run

Reduce GitHub Actions usage for the Delta Spark UT suite (per review feedback
on the pipeline PR):

* Per PR (velox_backend_x86.yml): a new `delta-changes` job runs the suite only
  when the PR touches high-signal Delta paths -- the Delta integration code
  (backends-velox/src-delta*), the gluten-delta module, or this pipeline's own
  files -- or carries the `run-delta-ci` opt-in label. Changes to general
  Velox/core/native code (touched on most PRs) skip it; this drops the per-PR
  trigger rate from ~60% to ~17% of recent commits.

* Nightly (delta_spark_ut.yml): a `schedule` (05:00 UTC) runs the full suite
  against the latest default branch, so regressions from the skipped-per-PR
  paths are still caught daily. It builds its own native lib (no caller) and
  uses fail_on_fixed=true, so baseline drift surfaces as a red nightly -- the
  signal to refresh known-failures.txt.

* Docs: README "When it runs" section documents the per-PR path gate, the
  opt-in label, and the nightly run.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Copilot-Session: 66a9e40f-3ac8-45be-8fee-a606a22fa098

* [CI] Fix stale Delta-job comment in velox_backend_x86.yml

The comment above the Delta Spark UT job still said it "runs on every trigger
like the other spark-test jobs", which contradicted the per-PR gating added by
the delta-changes job. Remove the stale block; keep the accurate gating comment
and add a concise native-lib-reuse note on the delta-spark-ut job.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Copilot-Session: 66a9e40f-3ac8-45be-8fee-a606a22fa098

* [CI] Baseline new ImplicitStreamingMergeCasting BIGINT->DECIMAL overflow failure

After the latest rebase, ImplicitStreamingMergeCastingSuite "Streaming MERGE
overflow sourceType: BIGINT, targetType: DECIMAL(7,2) followAnsiEnabled: false,
ansiEnabled: true, storeAssignmentPolicy: LEGACY" fails deterministically:
Velox raises `VeloxUserError INVALID_ARGUMENT: Cannot cast BIGINT
'9223372036854775807' to DECIMAL(7,2)` (rescaleInt, DecimalUtil.h) where vanilla
Spark handles the overflow per the ANSI/LEGACY policy. Its followAnsiEnabled:true
siblings are already baselined; add this variant so the gate stays green
(735 -> 736 entries, still sorted).

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Copilot-Session: 66a9e40f-3ac8-45be-8fee-a606a22fa098

* [CI] Skip already-applied Delta cherry-picks in setup-delta.sh

setup-delta.sh runs under `set -euo pipefail`, so cherry-picking a Delta
FileSourceScanLike fix that the pinned DELTA_REF already contains exits
non-zero (empty/conflicting patch) and aborts the whole setup. This is a
latent break the moment DELTA_REF is bumped past delta-io/delta#7104/#7105.

Make cherry_pick_delta_fix attempt the cherry-pick and, on failure, recover
only the paths that fix touches (git diff-tree -> per-file reset + checkout,
then cherry-pick --quit) and continue -- leaving the DeltaSQLCommandTest
patch and bundle jar intact. It is self-correcting: a genuinely missing fix
resurfaces as gate regressions instead of a hard abort.

Ancestry can't distinguish "already contained" from a real conflict here
because the Delta clone is shallow (depth 1) and merge-base --is-ancestor
can't see past the graft; a reverse-apply-check is fragile when the newer
ref carries the fix plus adjacent edits. Recover-on-failure avoids both.

Generated-by: GitHub Copilot CLI
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Copilot-Session: 66a9e40f-3ac8-45be-8fee-a606a22fa098

* [CI] Track skipped Delta tests separately and make the clone step idempotent

Two fixes from review feedback on the Delta gate.

1. compare-test-results.py treated a skipped test as "not seen this run", so a
   baseline entry that merely got skipped was reported under "Stale baseline
   entries (suite/test gone)" and silently dropped from the regenerated
   baseline -- only to return as a regression the next time it executed.

   Each shard now writes a --skipped-out list alongside --ran-out, and the
   aggregate job reports those entries as "Skipped this run" instead of stale
   and carries them over into the regenerated baseline. Genuinely removed tests
   appear in neither list and are still reported as stale.

   Skips are deliberately NOT folded into --ran-out: "now-passing" is derived
   from `ran - failed`, so counting a skip as a run would report it as fixed
   and, under fail_on_fixed, demand its removal from the baseline. Missing
   skipped-*.txt (older artifacts) degrades to the previous behaviour, and the
   shard-completeness guard still keys off failures-*/ran-* only.

2. setup-delta.sh could not be re-run over an existing DELTA_DIR: `git remote
   add origin` exits 3 when origin already exists, which aborts the script
   under `set -euo pipefail` and forces manual cleanup. Drop the remote first
   and force the checkout so a partial previous run is recovered rather than
   fatal. No `rm -rf` is reintroduced; the bundle jar and source patches are
   applied after this block, so nothing worth keeping is discarded.

Generated-by: GitHub Copilot CLI
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Copilot-Session: 66a9e40f-3ac8-45be-8fee-a606a22fa098

* [CI] Report invalid flaky-error regexes clearly and tighten the DV pattern

Follow-ups from review feedback on the Delta gate.

- load_patterns() compiled each line of flaky-error-patterns.txt with no error
  handling, so a typo in that hand-edited file surfaced as a bare re.error
  traceback naming neither the file nor the offending line. Raise
  BadPatternError with path, line number and pattern, and exit 2 from the gate:
  "<path>:4: invalid regex '[unclosed': unterminated character set at position 0".

- The negative-row-index signature used `-?\d+`, which would also match a
  positive index. The native check is `value >= 0`, so the reported index is
  always negative; match the sign explicitly, per this file's own rule of
  binding patterns to the exact error string so a quarantine can't swallow an
  unrelated failure.

- Note in the DeletionVectorsSuite block that the sed depends on the clone
  step's `checkout -f`. The sed appends after the test-declaration line, so
  without that per-run reset a re-run injects duplicate `fail` lines and trips
  the INJECTED != 2 check (verified: 2, 4, 6 without `-f`; 2, 2, 2 with it).

Generated-by: GitHub Copilot CLI
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Copilot-Session: 66a9e40f-3ac8-45be-8fee-a606a22fa098

* [CI] Make the Delta path gate fail open when git diff fails

The delta-changes job intends to run the Delta suite whenever it cannot
determine what the PR touched -- "never silently skip coverage". But the
detection itself was fail-closed:

  if git diff --name-only "$BASE" "$HEAD_SHA" | grep -Eq '<delta paths>'; then

When git diff fails (missing objects after a force-push race, an unfetched
fork head, or the `git merge-base ... || echo "$BASE_SHA"` fallback above
handing it an unusable sha), grep gets empty input and exits 1 -- exactly as
it does for "no Delta paths changed" -- so the else branch set run_delta=false
and the suite was skipped. `set -euo pipefail` does not help here: a command
used as an `if` condition is allowed to fail.

Capture the diff first and treat a git failure as an explicit fail-open, then
match against the captured output.

Verified over 8 scenarios: bad base+head, valid base + bad head and empty shas
now all yield run_delta=true, while gluten-delta/, backends-velox/src-delta40,
docs-only, cpp/velox-only and no-change diffs are unchanged. The previous
logic fails the first three.

Generated-by: GitHub Copilot CLI
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Copilot-Session: 66a9e40f-3ac8-45be-8fee-a606a22fa098

* [CI] Baseline new ImplicitMergeCasting BIGINT->DECIMAL overflow failure

Shard 2 of run 30324192115 reports one regression:

  ImplicitMergeCastingSuite#MERGE overflow in WHEN MATCHED THEN UPDATE SET
  t.value = s.value sourceType: BIGINT, targetType: DECIMAL(7,2)
  followAnsiEnabled: false, ansiEnabled: true, storeAssignmentPolicy: LEGACY

Under storeAssignmentPolicy LEGACY the overflowing value is expected to be
stored without error, but the offloaded Velox cast raises instead:

  VeloxUserError INVALID_ARGUMENT
  Reason: Cannot cast BIGINT '9223372036854775807' to DECIMAL(7, 2)
  Function: rescaleInt  File: velox/type/DecimalUtil.h:218

This is the non-streaming sibling of the ImplicitStreamingMergeCastingSuite
case for the same BIGINT -> DECIMAL(7,2) / LEGACY combination that was
baselined earlier; both surfaced after the rebase. It went unnoticed one run
longer because shard 2 of run 30186097028 died on a Docker pull before the
gate ran, so the test never executed there.

Baselined rather than quarantined in flaky-tests.txt: it failed in both runs
where it actually executed, with an identical error, and the failure is a pure
expression-evaluation overflow with no dependence on scheduling or runtime
plan -- unlike the DV bitmap row-index bug that flaky-error-patterns.txt
covers. Keeping it in the baseline means the gate will tell us to remove it
once the LEGACY cast semantics are fixed, which a flaky entry would not.

736 -> 737 entries.

Generated-by: GitHub Copilot CLI
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Copilot-Session: 66a9e40f-3ac8-45be-8fee-a606a22fa098

* [CI] Make the Delta hang-watchdog done marker shard-specific

run-delta-tests.sh is parameterized by SHARD_ID and already scopes its sbt log
and watchdog kill marker per shard, but the marker that tells the watchdog sbt
has finished was the global /tmp/sbt-done. With a shared /tmp -- parallel local
runs of several shards -- the first shard to finish creates it and every other
shard's watchdog exits its wait loop early, silently losing hang detection for
the shards still running.

Use /tmp/sbt-done-shard-${SHARD_ID} via a named SBT_DONE_MARKER, matching the
existing SBT_LOG / WATCHDOG_KILL_MARKER convention.

Verified with a reproduction of the arm/disarm handshake, running a fast shard
alongside a slow one: with the global marker the slow shard's watchdog was
disarmed after 3 ticks instead of its full 10; with the shard-scoped marker it
stays armed for the whole run. CI is unaffected (each shard is its own
container), so this only matters for local parallel runs.

Generated-by: GitHub Copilot CLI
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Copilot-Session: 66a9e40f-3ac8-45be-8fee-a606a22fa098

* [CI] Run the Delta Spark UT as its own workflow

The Delta suite ran as a reusable workflow called by velox_backend_x86.yml, so
its own pipeline files had to be in that workflow's `paths:` filter for a change
to them to be tested. The side effect was that a Delta-CI-only change pulled in
the entire Velox matrix: on this branch's last run, a commit touching a single
shell script produced 65 jobs and 2319 runner-minutes, of which 1750 (75%) were
TPC-H/DS and Spark UT jobs that the change could not affect.

Make delta_spark_ut.yml standalone. It already built its own native library for
the workflow_dispatch/schedule path, so this mostly means dropping
`workflow_call` and the `native_lib_artifact` input, adding a `pull_request`
trigger with a Delta `paths:` filter, and deleting the `delta-changes` gate and
`delta-spark-ut` jobs from velox_backend_x86.yml along with the Delta entries in
its `paths:`.

The `paths:` filter now IS the per-PR gate, evaluated before the run is created,
so an unrelated PR costs nothing at all -- this replaces the ~60-line
`delta-changes` script. That also drops the `run-delta-ci` label opt-in, which
cannot be expressed as a path filter: the label does not exist on apache/gluten
and so was never functional, and `workflow_dispatch` covers the same need (a
Velox/core author can run the suite against a branch, including on a fork).

Tradeoff: `gluten-delta/**` and `backends-velox/src-delta*/**` still match
velox_backend_x86.yml's filter -- it builds with -Pdelta -- so a change there now
runs both workflows and builds the native library twice (~10 min) instead of
sharing it. That is small next to the case above, and sharing it again is a
separate change.

Two conditions had to move with the split:

- `concurrency:` is now set. It was deliberately absent because, as a reusable
  workflow, `github.workflow` resolved to the caller and the group would have
  collided with the caller's own. Standalone it is needed, or every push to a
  Delta PR stacks another full run.
- delta-spark-aggregate no longer uses `always()`. That evaluates true while a
  run is being cancelled, so with `cancel-in-progress` a superseded push would
  cancel the shards but still start the aggregation, which then fails on finding
  no gate lists; the same happened when a failed bundle build skipped the shards.
  `!cancelled() && needs.delta-spark-test.result != 'skipped'` keeps the intended
  behaviour of publishing a baseline when shards go red, without the spurious
  second failure.

Also fix FAIL_ON_FIXED, which the split would otherwise have silently flipped:
`inputs` now only exists for workflow_dispatch, so the old expression resolved to
false on pull_request and schedule, where the removed workflow_call input had
supplied `default: true`. It now defaults to true for those events and honours
the input only on workflow_dispatch.

Generated-by: GitHub Copilot CLI
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Copilot-Session: 66a9e40f-3ac8-45be-8fee-a606a22fa098

* [CI] Split the Delta test matrix into 8 shards

Each Delta shard was taking ~2.5 hours, which sets the wall clock for the whole
pipeline. Breaking down one shard of a 4-shard run: 9.1 min of fixed setup
(clone Delta, apply the patches, compile the test sources) and 134.5 min actually
running tests, so ~91% of a shard is work that shards away.

Doubling to 8 shards therefore takes the slowest shard from ~148 min to ~74 min
for about +7% total runner-minutes, since the fixed setup is paid 8 times instead
of 4. This also matches delta-io/delta's own spark_test.yaml, which runs these
same suites with NUM_SHARDS: 8, shard: [0..7] and TEST_PARALLELISM_COUNT=4.

Memory is unaffected: each shard is a separate job, so it is still 4 forked test
JVMs plus the sbt launcher against the runner's ~16G, and TEST_PARALLELISM_COUNT
stays at 4.

No baseline regeneration is needed. The gate compares (suite, test) sets against
known-failures.txt -- regressions as `failed - baseline`, now-passing as
`baseline & passed`, stale as `baseline - ran - skipped` -- so it does not depend
on how suites are distributed across shards.

Sharding further would hit a floor at the longest single suite, currently ~18 min
(DeleteSQLSQLPathBasedDVPredPushOffSuite); 8 shards stays well clear of it, and
with ~180 suites per shard there is enough granularity for the split to balance.

Generated-by: GitHub Copilot CLI
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Copilot-Session: 66a9e40f-3ac8-45be-8fee-a606a22fa098

* [CI] Drop two now-passing DeltaUpdateCatalogSuite tests from the baseline

Shard 3 of the first 8-shard run went red on the now-passing check, not on a
regression: DeltaUpdateCatalogSuite's "convert to delta with partitioning
change" and "partitioned convert to delta with schema change" are in the
baseline but now pass. The aggregate job agrees globally -- 0 regressions,
2 now-passing, 0 stale.

Both tests previously failed with

  IllegalStateException: TaskResourceRegistry is not initialized

from ColumnarCachedBatchSerializer -> RowToVeloxColumnarExec ->
Runtimes.contextInstance, i.e. a task-listener setup problem on the cached-batch
path, which depends on what else is running in the same forked test JVM.
Resharding from 4 to 8 changed how suites are grouped into forks, and the
interleaving that triggered it no longer occurs.

This is a baseline update rather than a flaky quarantine: both tests failed in
each of the last two 4-shard runs and have been in the baseline since it was
bootstrapped, so there is no evidence of nondeterminism within a given
configuration -- the outcome tracked the shard count.

Verified by re-running the gate against that run's own gate-list artifacts with
this baseline: all 8 shards report 0 regressions and 0 now-passing, and the
aggregate's "distinct failing tests" (735) now matches the baseline exactly.

Generated-by: GitHub Copilot CLI
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Copilot-Session: 66a9e40f-3ac8-45be-8fee-a606a22fa098

---------

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Co-authored-by: Copilot Autofix powered by AI <175728472+Copilot@users.noreply.github.com>
Copilot-Session: 66a9e40f-3ac8-45be-8fee-a606a22fa098
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

BUILD CORE works for Gluten Core INFRA VELOX

Projects

None yet

Development

Successfully merging this pull request may close these issues.

6 participants