Repository navigation
[GLUTEN-12569][CORE] Add enhanced DSv2 filter test suites for Spark 4.2 - #13164
akshaytayal wants to merge 3 commits into
Conversation
|
Run Gluten Clickhouse CI on x86 |
3f80e82 to
4868b37
Compare
|
Run Gluten Clickhouse CI on x86 |
|
|
||
| class GlutenDataSourceV2EnhancedPartitionFilterSuite | ||
| extends DataSourceV2EnhancedPartitionFilterSuite | ||
| with GlutenSQLTestsTrait {} |
There was a problem hiding this comment.
Adapt the two residual-filter assertions to mixed native execution
Target Location: GlutenDataSourceV2EnhancedPartitionFilterSuite.scala:21-23.
Problem: Enabling this inherited suite with Gluten also enables two case9 tests whose final assertions require a Spark FilterExec. Their retained string-equality filters can correctly offload to FilterExecTransformer, which is not a subclass of FilterExec. The in-memory scan falls back, but scan fallback does not force the parent filter to fall back. With the workflow's ANSI=false and native filters enabled, correct mixed execution can therefore fail these vanilla-only plan assertions.
Evidence:
class GlutenDataSourceV2EnhancedPartitionFilterSuite
extends DataSourceV2EnhancedPartitionFilterSuite
with GlutenSQLTestsTrait {}The inherited "case 9" and nested identity case9 both use executedPlan.exists(_.isInstanceOf[FilterExec]). Gluten's filter transformer explicitly preserves the residual condition above a fallback child. This is source analysis, not an observed failure of these tests: the current group3 job stops at the parent's dependency error first.
Suggested Fix: Prefer Gluten variants of these two tests that retain all answer/pushdown checks and accept the intended native residual filter as well as a vanilla fallback filter. If this wrapper is deliberately scoped to Spark-filter/plugin integration instead, make that scope explicit without disabling the plugin:
import org.apache.spark.SparkConf
class GlutenDataSourceV2EnhancedPartitionFilterSuite
extends DataSourceV2EnhancedPartitionFilterSuite
with GlutenSQLTestsTrait {
override def sparkConf: SparkConf =
super.sparkConf.set("spark.gluten.sql.columnar.filter", "false")
}Run the suite with SPARK_ANSI_SQL_MODE=false; a pass with ANSI fallback enabled would not exercise the problematic mixed-plan path.
4868b37 to
8785e37
Compare
|
Run Gluten Clickhouse CI on x86 |
8785e37 to
18d7835
Compare
|
Run Gluten Clickhouse CI on x86 |
…r UTs for Spark 4.2 GlutenDataSourceV2EnhancedPartitionFilterSuite 'case 9 ...' (+ nested variant): the test asserts a post-scan vanilla FilterExec remains, but Gluten replaces it (columnar transform) / the InMemoryEnhancedPartitionFilterBatchScan falls back, so the plan-shape assertion fails. CI-confirmed (apache apache#13164 group3). Disabled via VeloxTestSettings.exclude; the other enhanced suites (Delete, RuntimePartition) pass. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> Copilot-Session: 413738d2-8cf1-4c70-9331-6c5e4314f9fa
|
Run Gluten Clickhouse CI on x86 |
…uites (apache#13164) No code changes; re-validate spark42 UT CI on the current branch (rebased on the green apache#13163 base + the 3 enhanced-filter commits). Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> Copilot-Session: 413738d2-8cf1-4c70-9331-6c5e4314f9fa
|
Run Gluten Clickhouse CI on x86 |
…Spark 4.2 Wrap Spark 4.2's DataSourceV2EnhancedRuntimePartitionFilterSuite with the Gluten test trait and enable it for Velox. This adds 15 upstream tests for the iterative runtime partition-filter path used by BatchScanExec. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> Copilot-Session: 8ae7d6bd-f561-417a-a635-40fe24dc67a5 (cherry picked from commit ed12775)
…or Spark 4.2 Wrap Spark 4.2's DataSourceV2EnhancedPartitionFilterSuite and DataSourceV2EnhancedDeleteFilterSuite with the Gluten test trait and enable them for Velox, completing coverage of the three new enhanced DSv2 filter suites. This adds 37 upstream tests covering first/second-pass partition filter pushdown, untranslatable and nested partition predicates, post-scan filter retention, and delete filter handling. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> Copilot-Session: 8ae7d6bd-f561-417a-a635-40fe24dc67a5 (cherry picked from commit 6a84498)
…r UTs for Spark 4.2 GlutenDataSourceV2EnhancedPartitionFilterSuite 'case 9 ...' (+ nested variant): the test asserts a post-scan vanilla FilterExec remains, but Gluten replaces it (columnar transform) / the InMemoryEnhancedPartitionFilterBatchScan falls back, so the plan-shape assertion fails. CI-confirmed (apache apache#13164 group3). Disabled via VeloxTestSettings.exclude; the other enhanced suites (Delete, RuntimePartition) pass. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> Copilot-Session: 413738d2-8cf1-4c70-9331-6c5e4314f9fa
3d93fe7 to
a9a0290
Compare
|
Run Gluten Clickhouse CI on x86 |
What changes were proposed in this pull request?
Adds the enhanced DataSourceV2 filter test suites for Spark 4.2 in
gluten-ut/spark42, wrapping the upstream Apache Spark suites withGlutenSQLTestsTrait:GlutenDataSourceV2EnhancedRuntimePartitionFilterSuiteGlutenDataSourceV2EnhancedPartitionFilterSuiteGlutenDataSourceV2EnhancedDeleteFilterSuiteand registers them in
VeloxTestSettings.Disabled (CI-confirmed)
2 tests in
GlutenDataSourceV2EnhancedPartitionFilterSuiteare disabled viaVeloxTestSettings.excludebecause Gluten changes the plan shape they assert on:case 9: partition filter pushed but returned in first pass is not re-pushed second passnested identity partition: case 9 partition filter pushed but returned in first pass is not re-pushed second passRoot cause: the tests assert a post-scan vanilla
FilterExecremains, but Gluten replaces it with a columnarFilterExecTransformer/ theInMemoryEnhancedPartitionFilterBatchScanfalls back. The other two enhanced suites (Delete, RuntimePartition) pass. Follow-up to re-enable tracked with the DSv2 filter work (see #13181 umbrella).How was this patch tested?
spark42 UT CI (groups 1-3 + slow extended/slow-hive) is green on this branch, which is rebased on top of #13163 (so it also carries the Spark 4.2 UT module + all its CI-green fixes).
Notes
(GLUTEN-12569)
Related issue: #12569