Skip to content

fix: restore localReads assertion in AdaptiveQueryExecSuite diffs - #6382

Open
mizulun wants to merge 3 commits into
apache:mainfrom
mizulun:6122-v1
Open

mizulun wants to merge 3 commits into
apache:mainfrom
mizulun:6122-v1

Conversation

@mizulun

@mizulun mizulun commented Sep 29, 2026

Copy link
Copy Markdown
Contributor

Which issue does this PR close?

Closes #6122.

Rationale for this change

In AdaptiveQueryExecSuite, the test Reuse the default parallelism in local shuffle read had its assertion changed from Spark's assert(localReads.length == 2) to == 1 in the 3.5.9, 4.0.4 and 4.1.3 diffs. The next statement reads localReads(1), so the test cannot pass for any length. With one local read it throws IndexOutOfBoundsException, and with two the assertion fails.

CI stays green because the test carries IgnoreComet, so it only runs when Comet is disabled. That is the Spark-only baseline run. The 4.2.0 diff already keeps Spark's value, and this PR brings the other three diffs in line with it.

What changes are included in this PR?

  • dev/diffs/3.5.9.diff, dev/diffs/4.0.4.diff, dev/diffs/4.1.3.diff: drop the
    // Comet shuffle changes shuffle metrics comment and restore assert(localReads.length == 2).
    The IgnoreComet tag on the test stays.
  • Each diff was regenerated as described in spark-sql-tests.md: apply it to the Spark tag, edit the source, then run git diff. Apart from the removed hunk, the only changes are the line offsets of later hunks in AdaptiveQueryExecSuite.scala.

How are these changes tested?

For each Spark version, the diff was applied to its tag checkout and the test was run in both modes:

NOLINT_ON_COMPILE=true ENABLE_COMET=<false|true> build/sbt \
  "sql/testOnly org.apache.spark.sql.execution.adaptive.AdaptiveQueryExecSuite -- -z \"Reuse the default parallelism in local shuffle read\""
Spark ENABLE_COMET=false ENABLE_COMET=true
3.5.9 succeeded 1 ignored 1
4.0.4 succeeded 1 ignored 1
4.1.3 succeeded 1 ignored 1

The baseline run now exercises the restored assertion, and the Comet run still skips the test through the existing IgnoreComet tag.

@github-actions github-actions Bot added the bug Something isn't working label Sep 29, 2026

@sunchao sunchao left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Summary

  • Prior state and problem: Three Spark patches changed localReads.length to 1 while retaining access to localReads(1), making the Spark-only baseline test impossible to pass.
  • Design approach: Restore Spark’s assertion by removing the overriding hunk from the 3.5.9, 4.0.4 and 4.1.3 patches.
  • Correctness / compatibility analysis: Applied the base and head AQE patch sections to each matching Spark tag. All applied successfully and matched their declared blob hashes. Each restored test body matches upstream exactly, apart from the retained IgnoreComet tag.
  • Key design decisions: Preserve Comet’s existing test exclusion and reuse upstream behavior without adding branches or abstractions.
  • Implementation sketch: Remove the incorrect assertion and accompanying comment. All other changes are verified patch hashes and downstream offsets.
  • Behavioral changes worth calling out: The baseline test again expects two local reads. Comet-enabled runs still skip it. Production execution and performance are unchanged.
  • Suggested improvements: None meeting the reporting threshold. No introduced P1/P2 issues found within this review.

Reviewed the complete diff from 4ca04b155bed7abe9e0e267f6438c31cca2c278d to f39539fa12382817a472f30446f97a8c57ea2396. The PR is not a draft. Read the supplied snapshot, live discussion and linked issue #6122. No existing review concerns remain unresolved.

Routed skills: review-comet-pr and review-comet-shuffle-pr.

Exact-head CI: label passed. Comet CI and CodeQL report action_required. No build or Spark-suite verdict is available.

Validation limits: Patch application, source comparison, offset checks and git diff --check passed. Scala/Spark suites were not run locally because no prepared Spark/Comet build was available. The author-reported runtime results were not independently reproduced.

@andygrove

Copy link
Copy Markdown
Member

dev/diffs/3.4.3.diff has the same broken edit, and this PR doesn't touch it. In Reuse the default parallelism in local shuffle read it changes assert(localReads.length == 2) to == 1 (around line 1681 of the diff on main), and localReads(1) still follows two lines later. I missed it when I filed #6122 and only listed three diffs. Could we fix 3.4.3 in this PR as well, using the same apply, edit and git diff steps from spark-sql-tests.md against v3.4.3? Then Closes #6122 won't leave a broken diff behind.

@mizulun

mizulun commented Oct 1, 2026

Copy link
Copy Markdown
Contributor Author

Thanks for catching this, @andygrove ! I’ve updated the 3.4.3 diff and pushed the fix.

@andygrove andygrove left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks @mizulun

@andygrove

Copy link
Copy Markdown
Member

@mizulun could you fix conflicts?

@mizulun

mizulun commented Oct 9, 2026

Copy link
Copy Markdown
Contributor Author

Sure, I’ll fix the conflicts as soon as possible. Thanks for the reminder!

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

bug Something isn't working

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Three dev/diffs weaken a local-shuffle-read assertion into a contradiction, breaking the Spark-only baseline

3 participants