Skip to content

[Spark 4.2 UT] Spark 4.2 Storage-Partitioned Join (SPJ) support in Gluten #13174

Description

@akshaytayal

Part of the Spark 4.2 unit-test enablement (GLUTEN-12569, PR #13163). This issue tracks re-enabling one root-cause group of UTs that are currently disabled to keep the gluten-ut/spark42 CI green.

Root cause

Spark 4.2 Storage-Partitioned Join (SPJ) support in Gluten

How it's disabled

.exclude(...) in gluten-ut/spark42/src/test/scala/org/apache/gluten/utils/velox/VeloxTestSettings.scala (under enableSuite[GlutenKeyGroupedPartitioningSuite]).

These suites honor the Gluten exclude registry (BackendTestSettings.shouldRun), so a .exclude("<test>") line skips the test.

Files updated

  • gluten-ut/spark42/src/test/scala/org/apache/gluten/utils/velox/VeloxTestSettings.scala

Disabled tests (29)

  • GlutenKeyGroupedPartitioningSuite :: Gluten - SPARK-42038: partially clustered: full outer join is not applicable
  • GlutenKeyGroupedPartitioningSuite :: Gluten - SPARK-42038: partially clustered: left outer join
  • GlutenKeyGroupedPartitioningSuite :: Gluten - SPARK-42038: partially clustered: right outer join
  • GlutenKeyGroupedPartitioningSuite :: Gluten - SPARK-42038: partially clustered: with different partition keys and both sides partially clustered
  • GlutenKeyGroupedPartitioningSuite :: Gluten - SPARK-42038: partially clustered: with different partition keys and missing keys on left-hand side
  • GlutenKeyGroupedPartitioningSuite :: Gluten - SPARK-42038: partially clustered: with different partition keys and missing keys on right-hand side
  • GlutenKeyGroupedPartitioningSuite :: Gluten - SPARK-42038: partially clustered: with same partition keys and both sides partially clustered
  • GlutenKeyGroupedPartitioningSuite :: Gluten - SPARK-42038: partially clustered: with same partition keys and one side fully clustered
  • GlutenKeyGroupedPartitioningSuite :: Gluten - SPARK-44647: SPJ: test join key is subset of cluster key with push values and partially-clustered
  • GlutenKeyGroupedPartitioningSuite :: Gluten - SPARK-44647: test join key is the second partition key and a transform
  • GlutenKeyGroupedPartitioningSuite :: Gluten - SPARK-47094: Compatible buckets does not support SPJ with push-down values or partially-clustered
  • GlutenKeyGroupedPartitioningSuite :: Gluten - SPARK-53322: checkpointed scans aren't used for SPJ
  • GlutenKeyGroupedPartitioningSuite :: Gluten - SPARK-53322: checkpointed scans can be shuffled by children on SPJ
  • GlutenKeyGroupedPartitioningSuite :: Gluten - SPARK-53322: checkpointed scans can't shuffle other children on SPJ
  • GlutenKeyGroupedPartitioningSuite :: SPARK-48065: SPJ: allowJoinKeysSubsetOfPartitionKeys is too strict
  • GlutenKeyGroupedPartitioningSuite :: SPARK-55535: Multi table join granular partition grouping
  • GlutenKeyGroupedPartitioningSuite :: SPARK-55535: Multi table join partial clustering
  • GlutenKeyGroupedPartitioningSuite :: SPARK-55715: Custom metrics of sorted-merge coalesced partitions
  • GlutenKeyGroupedPartitioningSuite :: SPARK-55715: preserve outputOrdering when coalescing partitions with sorted merge
  • GlutenKeyGroupedPartitioningSuite :: SPARK-55715: preserve outputOrdering when coalescing transform-partitioned splits
  • GlutenKeyGroupedPartitioningSuite :: SPARK-55848: Window dedup after SPJ with partial clustering
  • GlutenKeyGroupedPartitioningSuite :: SPARK-55848: checkpointed partially-clustered join with dedup
  • GlutenKeyGroupedPartitioningSuite :: SPARK-55848: dropDuplicates after SPJ with partial clustering
  • GlutenKeyGroupedPartitioningSuite :: SPARK-56241: GroupPartitionsExec coalescing derives ordering from key expressions, no pre-join SortExec needed before SortMergeJoin
  • GlutenKeyGroupedPartitioningSuite :: SPARK-56241: GroupPartitionsExec non-coalescing passes through child ordering, no pre-join SortExec needed before SortMergeJoin
  • GlutenKeyGroupedPartitioningSuite :: SPARK-56549: k-way merge enabled only when parent requires ordering
  • GlutenKeyGroupedPartitioningSuite :: partitioned join: exact distribution (same number of buckets) from both sides
  • GlutenKeyGroupedPartitioningSuite :: partitioned join: join with two partition keys and matching & sorted partitions
  • GlutenKeyGroupedPartitioningSuite :: partitioned join: join with two partition keys and unsorted partitions

How to re-enable

Implement the underlying Gluten/Velox support, then remove the corresponding .exclude(...) / allowlist / @Ignore / ignore(...) entries and confirm the tests pass in the spark42 UT CI.

Related: PR #13163, GLUTEN-12569.

Umbrella: #13181

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions