Repository navigation
fix: enable FIRST/LAST partial merge - #5041
Merged
comphead merged 5 commits intoSep 9, 2026
Merged
Conversation
peterxcli
marked this pull request as ready for review
September 5, 2026 19:28
peterxcli
force-pushed
the
codex/issue-4131-last-value-partial-merge
branch
from
September 5, 2026 20:19
ea2e9b7 to
11fdb9a
Compare
Member
Author
|
cc @comphead ptal, thanks! |
comphead
approved these changes
Sep 8, 2026
comphead
left a comment
Contributor
There was a problem hiding this comment.
Thanks @peterxcli it looks good to me, btw can we try to move impacted tests from suites to sql files if that possible
Contributor
|
Hm, this PR ran incomplete CI |
comphead
self-requested a review
September 8, 2026 23:09
Member
Author
done, thanks for the suggestion! |
andygrove
added a commit
that referenced
this pull request
Oct 3, 2026
…ck (#5421) (#6422) * fix: [branch-1.0] revert unsafe partial aggregates after final fallback (#5421) Backport of #5421 to branch-1.0. Adaptations for branch-1.0, which does not have the Celeborn shuffle planning from #5537: - CometExecRule: keep restoreSparkPartial for the new repair pass and drop preserveSparkAggregateBuffers. The early tagging pass keeps branch-1.0's Final-only consumer check and switches to the renamed predicate. - RevertNativeForTransitionHeavyStages: drop the hunk. It edits hasUnsafeMixedAggregateAtStageBoundary, which branch-1.0 does not have. - CometCelebornShufflePlanningSuite: not on branch-1.0, dropped. - Tests: take the #5421 tests and the imports they use. The FIRST/LAST percentile test in the conflict comes from #5041, which is not on branch-1.0. (cherry picked from commit 6065705) * fix: [branch-1.0] stop transition reversion from splitting unsafe aggregates RevertNativeForTransitionHeavyStages runs after the #5421 repair pass, and on branch-1.0 it can still revert the stage that holds a native Final while the native Partial in the stage below stays native. With spark.comet.exec.transitionRevert.enabled=true this brings back the #5419 failures: AVG returns null when a scan partition is empty, and a Spark Final cannot read the native collect_list or percentile buffer. Port the hasUnsafeMixedAggregateAtStageBoundary guard from main, where #5537 added it and #5421 switched it to the directional predicates. The rule now matches main as of #5421. Also port the #5537 tests for the guard, including its switch from COUNT to MIN and MAX in two existing tests, because the guard keeps both COUNT stages native. Add a test for the AVG and collect_list cases, which the #5537 tests do not cover. --------- Co-authored-by: Chao Sun <sunchao@apache.org>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Which issue does this PR close?
Closes #4131.
Rationale for this change
DataFusion LAST_VALUE partial-state merging filtered unset states but selected element 0 from the final state array instead of its last element. This could make Comet produce a different result from Spark, so Comet temporarily fell back for FIRST/LAST in PartialMerge mode.
Upstream
mainuses DataFusion 55.0.0 and Arrow/Parquet 59.2.0 through merged #5262. The dependency upgrade is no longer a blocker. The root fix, apache/datafusion#23905, is included in DataFusion 55.0.0.What changes are included in this PR?
How are these changes tested?
The following results are from the backport-based draft, not a DataFusion 55.0.0 validation run. Rerun them after the rebase and dependency-pin cleanup.
cargo build --lockedJAVA_HOME=/Users/lixucheng/.sdkman/candidates/java/17.0.14-zulu make coreJAVA_HOME=/Users/lixucheng/.sdkman/candidates/java/17.0.14-zulu DYLD_LIBRARY_PATH=/Users/lixucheng/.sdkman/candidates/java/17.0.14-zulu/lib/server ./mvnw test -Dtest=none -Dsuites="org.apache.comet.CometSqlFileTestSuite partial_merge" -Dscalastyle.skip=trueThe focused SQL suite passed both Parquet dictionary configurations: 2 tests run, 2 succeeded.