Skip to content

[GLUTEN-12569][CORE] Add enhanced DSv2 filter test suites for Spark 4.2 - #13164

Draft
akshaytayal wants to merge 3 commits into
apache:mainfrom
akshaytayal:spark42-ut-ext
Draft

akshaytayal wants to merge 3 commits into
apache:mainfrom
akshaytayal:spark42-ut-ext

Conversation

@akshaytayal

@akshaytayal akshaytayal commented Sep 30, 2026 •

Copy link
Copy Markdown
Contributor

What changes were proposed in this pull request?

Adds the enhanced DataSourceV2 filter test suites for Spark 4.2 in gluten-ut/spark42, wrapping the upstream Apache Spark suites with GlutenSQLTestsTrait:

  • GlutenDataSourceV2EnhancedRuntimePartitionFilterSuite
  • GlutenDataSourceV2EnhancedPartitionFilterSuite
  • GlutenDataSourceV2EnhancedDeleteFilterSuite

and registers them in VeloxTestSettings.

Disabled (CI-confirmed)

2 tests in GlutenDataSourceV2EnhancedPartitionFilterSuite are disabled via VeloxTestSettings.exclude because Gluten changes the plan shape they assert on:

  • case 9: partition filter pushed but returned in first pass is not re-pushed second pass
  • nested identity partition: case 9 partition filter pushed but returned in first pass is not re-pushed second pass

Root cause: the tests assert a post-scan vanilla FilterExec remains, but Gluten replaces it with a columnar FilterExecTransformer / the InMemoryEnhancedPartitionFilterBatchScan falls back. The other two enhanced suites (Delete, RuntimePartition) pass. Follow-up to re-enable tracked with the DSv2 filter work (see #13181 umbrella).

How was this patch tested?

spark42 UT CI (groups 1-3 + slow extended/slow-hive) is green on this branch, which is rebased on top of #13163 (so it also carries the Spark 4.2 UT module + all its CI-green fixes).

Notes

(GLUTEN-12569)

Related issue: #12569

@github-actions

Copy link
Copy Markdown

Run Gluten Clickhouse CI on x86

@github-actions

Copy link
Copy Markdown

Run Gluten Clickhouse CI on x86


class GlutenDataSourceV2EnhancedPartitionFilterSuite
extends DataSourceV2EnhancedPartitionFilterSuite
with GlutenSQLTestsTrait {}

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Adapt the two residual-filter assertions to mixed native execution

Target Location: GlutenDataSourceV2EnhancedPartitionFilterSuite.scala:21-23.

Problem: Enabling this inherited suite with Gluten also enables two case9 tests whose final assertions require a Spark FilterExec. Their retained string-equality filters can correctly offload to FilterExecTransformer, which is not a subclass of FilterExec. The in-memory scan falls back, but scan fallback does not force the parent filter to fall back. With the workflow's ANSI=false and native filters enabled, correct mixed execution can therefore fail these vanilla-only plan assertions.

Evidence:

class GlutenDataSourceV2EnhancedPartitionFilterSuite
  extends DataSourceV2EnhancedPartitionFilterSuite
  with GlutenSQLTestsTrait {}

The inherited "case 9" and nested identity case9 both use executedPlan.exists(_.isInstanceOf[FilterExec]). Gluten's filter transformer explicitly preserves the residual condition above a fallback child. This is source analysis, not an observed failure of these tests: the current group3 job stops at the parent's dependency error first.

Suggested Fix: Prefer Gluten variants of these two tests that retain all answer/pushdown checks and accept the intended native residual filter as well as a vanilla fallback filter. If this wrapper is deliberately scoped to Spark-filter/plugin integration instead, make that scope explicit without disabling the plugin:

import org.apache.spark.SparkConf

class GlutenDataSourceV2EnhancedPartitionFilterSuite
  extends DataSourceV2EnhancedPartitionFilterSuite
  with GlutenSQLTestsTrait {
  override def sparkConf: SparkConf =
    super.sparkConf.set("spark.gluten.sql.columnar.filter", "false")
}

Run the suite with SPARK_ANSI_SQL_MODE=false; a pass with ANSI fallback enabled would not exercise the problematic mixed-plan path.

@github-actions

github-actions Bot commented Oct 1, 2026

Copy link
Copy Markdown

Run Gluten Clickhouse CI on x86

@github-actions

github-actions Bot commented Oct 2, 2026

Copy link
Copy Markdown

Run Gluten Clickhouse CI on x86

akshaytayal added a commit to akshaytayal/gluten that referenced this pull request Oct 2, 2026
…r UTs for Spark 4.2

GlutenDataSourceV2EnhancedPartitionFilterSuite 'case 9 ...' (+ nested variant):
the test asserts a post-scan vanilla FilterExec remains, but Gluten replaces it
(columnar transform) / the InMemoryEnhancedPartitionFilterBatchScan falls back,
so the plan-shape assertion fails. CI-confirmed (apache apache#13164 group3). Disabled
via VeloxTestSettings.exclude; the other enhanced suites (Delete, RuntimePartition)
pass.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Copilot-Session: 413738d2-8cf1-4c70-9331-6c5e4314f9fa
@github-actions

github-actions Bot commented Oct 2, 2026

Copy link
Copy Markdown

Run Gluten Clickhouse CI on x86

akshaytayal added a commit to akshaytayal/gluten that referenced this pull request Oct 2, 2026
…uites (apache#13164)

No code changes; re-validate spark42 UT CI on the current branch (rebased on the
green apache#13163 base + the 3 enhanced-filter commits).

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Copilot-Session: 413738d2-8cf1-4c70-9331-6c5e4314f9fa
@github-actions

github-actions Bot commented Oct 2, 2026

Copy link
Copy Markdown

Run Gluten Clickhouse CI on x86

MANOJ RAGUPATHY and others added 3 commits October 10, 2026 09:46
…Spark 4.2

Wrap Spark 4.2's DataSourceV2EnhancedRuntimePartitionFilterSuite with the
Gluten test trait and enable it for Velox. This adds 15 upstream tests for
the iterative runtime partition-filter path used by BatchScanExec.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Copilot-Session: 8ae7d6bd-f561-417a-a635-40fe24dc67a5
(cherry picked from commit ed12775)
…or Spark 4.2

Wrap Spark 4.2's DataSourceV2EnhancedPartitionFilterSuite and
DataSourceV2EnhancedDeleteFilterSuite with the Gluten test trait and enable
them for Velox, completing coverage of the three new enhanced DSv2 filter
suites.

This adds 37 upstream tests covering first/second-pass partition filter
pushdown, untranslatable and nested partition predicates, post-scan filter
retention, and delete filter handling.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Copilot-Session: 8ae7d6bd-f561-417a-a635-40fe24dc67a5
(cherry picked from commit 6a84498)
…r UTs for Spark 4.2

GlutenDataSourceV2EnhancedPartitionFilterSuite 'case 9 ...' (+ nested variant):
the test asserts a post-scan vanilla FilterExec remains, but Gluten replaces it
(columnar transform) / the InMemoryEnhancedPartitionFilterBatchScan falls back,
so the plan-shape assertion fails. CI-confirmed (apache apache#13164 group3). Disabled
via VeloxTestSettings.exclude; the other enhanced suites (Delete, RuntimePartition)
pass.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Copilot-Session: 413738d2-8cf1-4c70-9331-6c5e4314f9fa
@github-actions

Copy link
Copy Markdown

Run Gluten Clickhouse CI on x86

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

CORE works for Gluten Core

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants