Repository navigation
Conversation
MAX_BY and MIN_BY accept a StringType value with the default UTF8_BINARY collation; the ordering stays limited to fixed-length types. The native MaxMinBy accumulators already store the value as Arrow row bytes and only compare the ordering, so no native change is needed. A string value makes Spark plan a SortAggregate, which is now converted when its other aggregates are order-insensitive. MAX_BY and MIN_BY depend on the input order only among rows tied on the ordering, where Spark is non-deterministic too and the native path already documents that it may pick a different row. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Collaborator
Author
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Depends on #12. The base branch is #12's branch; merge this right after #12.
What
MAX_BY/MIN_BYaccept aStringTypevalue with the default UTF8_BINARY collation. The ordering is still limited to fixed-length types; a string ordering, a collated value and other variable-length types stay in Spark.MaxMinBy(both the single and the groups accumulator) stores the value as Arrow row bytes and only compares the ordering. Its native unit tests already use a Utf8 value.MaxMinBynow counts as such.Semantics vs Spark
old > new, or<formin_by), and its merge keeps the incoming buffer. Which tied row wins therefore depends on input and merge order, and Spark's documentation callsmax_bynon-deterministic. The native accumulator also keeps the later row of a batch, but it may still pick a different tied row than Spark. This is the same difference the existing native path for fixed-length values already documents ingetCompatibleNotes. Compares should not expect equality on rows tied on the ordering.max_by/min_bydepend on the input order only among tied rows. FIRST/LAST/ANY_VALUE stay in Spark per the feat: native MIN/MAX over strings and SortAggregate as a native hash aggregate #12 fix.Tests
CometAggregateSuite:max_by/min_byover strings with nulls, an empty string, non-ASCII text and null orderings; grouped and ungrouped; AQE on and off (partial/final). The SortAggregate is gone and results equal Spark.max_by.sql/min_by.sql: the string-value fallback expectations now run natively, plus new string-value queries. The string-ordering fallbacks are unchanged.🤖 Generated with Claude Code