CometAggregateSuite |
63 |
grouped aggregate metrics are forwarded without fabricating global metrics (ignored) |
keep |
|
CometAggregateSuite |
105 |
range sampling does not report grouped aggregate metrics (ignored) |
keep |
|
CometAggregateSuite |
157 |
collect_list over struct with non-nullable fields |
convert |
aggregate/collect_list.sql |
CometAggregateSuite |
191 |
grouped collect_list/collect_set over nulls, duplicates and several batches |
convert |
aggregate/collect_grouped_batches.sql (new) |
CometAggregateSuite |
215 |
collect_list/collect_set over typed nested NULL (grouped=$grouped) [instances: grouped=false, grouped=true] |
convert |
aggregate/collect_list.sql |
CometAggregateSuite |
240 |
collect_list/collect_set combined with distinct aggregate falls back safely |
convert |
aggregate/collect_distinct_fallback.sql (new) |
CometAggregateSuite |
276 |
min/max floating point with negative zero |
convert |
aggregate/min_max_strict_floating_point.sql (new); aggregate/group_by_floating_point.sql (new) |
CometAggregateSuite |
309 |
avg decimal |
convert |
aggregate/sum_avg_decimal.sql (new) |
CometAggregateSuite |
340 |
approx_percentile with distinct aggregate does not split across Comet and Spark |
convert |
aggregate/approx_percentile_distinct.sql (new) |
CometAggregateSuite |
359 |
disabled FIRST/LAST preserves percentile buffers with local Comet shuffle |
keep |
|
CometAggregateSuite |
388 |
decimal AVG falls back across a Spark shuffle (AQE=$adaptive) [instances: AQE=false, AQE=true] |
keep |
|
CometAggregateSuite |
455 |
COUNT and AVG fall back together across a Spark shuffle (AQE=$adaptive) [instances: AQE=false, AQE=true] |
keep |
|
CometAggregateSuite |
512 |
COUNT preserves safe native partials across a Spark shuffle (AQE=$adaptive) [instances: AQE=false, AQE=true] |
keep |
|
CometAggregateSuite |
540 |
$fn falls back when enabled native shuffle is ineligible (AQE=$adaptive) [instances: collect_list/collect_set x AQE=false/true] |
keep |
|
CometAggregateSuite |
580 |
$fn preserves aggregate buffers with an unsupported array hash key (AQE=$adaptive) [instances: percentile/collect_list/sum x AQE=false/true] |
keep |
|
CometAggregateSuite |
671 |
stddev_pop should return NaN for some cases |
convert |
aggregate/stddev.sql |
CometAggregateSuite |
686 |
count with aggregation filter |
convert |
aggregate/distinct.sql (new) |
CometAggregateSuite |
699 |
multiple column distinct count |
convert |
aggregate/distinct.sql (new) |
CometAggregateSuite |
718 |
Only trigger Comet Final aggregation on Comet partial aggregation |
convert |
aggregate/mixed_engine_spark_partial.sql (new) |
CometAggregateSuite |
731 |
Average expression in Comet Final should handle all null inputs from partial Spark aggregation |
convert |
aggregate/mixed_engine_spark_partial.sql (new) |
CometAggregateSuite |
746 |
decimal SUM partial stays in Spark when a later input cancels precision overflow |
keep |
|
CometAggregateSuite |
786 |
mixed engine sum/avg falls back when Spark Final would consume native AVG |
keep |
|
CometAggregateSuite |
800 |
mixed engine sum/avg: Spark partial + Comet final matches Spark |
convert |
aggregate/mixed_engine_partial_disabled.sql (new) |
CometAggregateSuite |
823 |
mixed engine collect_list: $name matches Spark [instances: 'Comet partial + Spark final', 'Spark partial + Comet final'] |
keep |
|
CometAggregateSuite |
840 |
Aggregation without aggregate expressions should use correct result expressions |
convert |
aggregate/distinct.sql (new) |
CometAggregateSuite |
856 |
Final aggregation should not bind to the input of partial aggregation |
convert |
aggregate/group_by_numeric_keys.sql (new) |
CometAggregateSuite |
874 |
Ensure traversed operators during finding first partial aggregation are all native |
convert |
aggregate/correlated_subquery_aggregate.sql (new) |
CometAggregateSuite |
911 |
SUM decimal supports emit.first |
blocked: #6618 |
aggregate/sum_avg_decimal_sorted_input.sql (new) |
CometAggregateSuite |
930 |
AVG decimal supports emit.first |
blocked: #6618 |
aggregate/sum_avg_decimal_sorted_input.sql (new) |
CometAggregateSuite |
949 |
Fix NPE in partial decimal sum |
convert |
aggregate/sum_avg_decimal.sql (new) |
CometAggregateSuite |
966 |
fix: Decimal Average should not enable native final aggregation |
convert |
aggregate/sum_avg_decimal.sql (new) |
CometAggregateSuite |
982 |
trivial case |
convert |
aggregate/group_by_numeric_keys.sql (new) |
CometAggregateSuite |
999 |
avg |
convert |
aggregate/group_by_numeric_keys.sql (new) |
CometAggregateSuite |
1012 |
count, avg with null |
convert |
aggregate/group_by_numeric_keys.sql (new) |
CometAggregateSuite |
1028 |
SUM/AVG non-decimal overflow |
convert |
aggregate/sum.sql |
CometAggregateSuite |
1037 |
simple SUM, COUNT, MIN, MAX, AVG with non-distinct group keys |
convert |
aggregate/group_by_numeric_keys.sql (new) |
CometAggregateSuite |
1050 |
group-by on variable length types |
split (rest: plan shape stays in Scala) |
aggregate/avg.sql |
CometAggregateSuite |
1075 |
grouping on struct containing map should fallback to Spark |
convert |
aggregate/group_by_map_fallback.sql (new) |
CometAggregateSuite |
1102 |
simple SUM, COUNT, MIN, MAX, AVG with non-distinct + null group keys |
convert |
aggregate/group_by_numeric_keys.sql (new) |
CometAggregateSuite |
1123 |
simple SUM, COUNT, MIN, MAX, AVG with null aggregates |
convert |
aggregate/group_by_numeric_keys.sql (new) |
CometAggregateSuite |
1147 |
simple SUM, MIN, MAX, AVG with all nulls |
convert |
aggregate/group_by_numeric_keys.sql (new) |
CometAggregateSuite |
1175 |
SUM, COUNT, MIN, MAX, AVG on float & double |
convert |
aggregate/group_by_floating_point.sql (new) |
CometAggregateSuite |
1190 |
SUM, MIN, MAX, AVG for NaN, -0.0 and 0.0 |
convert |
aggregate/group_by_floating_point.sql (new) |
CometAggregateSuite |
1209 |
SUM/MIN/MAX/AVG on decimal |
convert |
aggregate/sum_avg_decimal.sql (new) |
CometAggregateSuite |
1223 |
multiple SUM/MIN/MAX/AVG on decimal and non-decimal |
convert |
aggregate/sum_avg_decimal.sql (new) |
CometAggregateSuite |
1242 |
SUM/AVG on decimal with different precisions |
convert |
aggregate/sum_avg_decimal.sql (new) |
CometAggregateSuite |
1263 |
SUM decimal with DF |
split (rest: plan shape stays in Scala) |
aggregate/sum_avg_decimal.sql (new) |
CometAggregateSuite |
1303 |
COUNT/MIN/MAX on date, timestamp |
convert |
aggregate/min_max_datetime.sql (new) |
CometAggregateSuite |
1320 |
single group-by column + aggregate column, multiple batches, no null |
convert |
aggregate/group_by_multi_batch.sql (new) |
CometAggregateSuite |
1347 |
partialMerge - cnt distinct + sum |
convert |
aggregate/partial_merge.sql |
CometAggregateSuite |
1388 |
partialMerge - distinct + non-distinct aggregates (Expand pattern) |
convert |
aggregate/partial_merge.sql |
CometAggregateSuite |
1402 |
multiple group-by columns + single aggregate column (first/last), with nulls |
convert |
aggregate/first_last_multi_batch.sql (new) |
CometAggregateSuite |
1434 |
multiple group-by columns + single aggregate column, with nulls |
convert |
aggregate/group_by_multi_batch.sql (new) |
CometAggregateSuite |
1467 |
string should be supported |
convert |
aggregate/sum.sql |
CometAggregateSuite |
1477 |
multiple group-by columns + multiple aggregate column (first/last), with nulls |
convert |
aggregate/first_last_multi_batch.sql (new) |
CometAggregateSuite |
1506 |
multiple group-by columns + multiple aggregate column, with nulls |
convert |
aggregate/group_by_multi_batch.sql (new) |
CometAggregateSuite |
1539 |
all types first/last, with nulls |
convert |
aggregate/group_by_all_types.sql (new) |
CometAggregateSuite |
1571 |
first/last with ignore null |
convert |
aggregate/first_last_multi_batch.sql (new) |
CometAggregateSuite |
1588 |
all types, with nulls |
convert |
aggregate/group_by_all_types.sql (new) |
CometAggregateSuite |
1619 |
test final count |
convert |
aggregate/group_by_numeric_keys.sql (new) |
CometAggregateSuite |
1636 |
test final min/max |
convert |
aggregate/group_by_numeric_keys.sql (new) |
CometAggregateSuite |
1657 |
test final min/max/count with result expressions |
split (rest: plan shape stays in Scala) |
aggregate/group_by_numeric_keys.sql (new) |
CometAggregateSuite |
1698 |
aggregate canonicalization preserves result expressions and equivalent reuse: $function, AQE=$adaptive [instances: COUNT()/false, COUNT(DISTINCT _2)/false, COUNT(DISTINCT _2) + SUM(_2)/false, CAST(SIZE(COLLECT_SET(_2)) AS BIGINT)/false, COUNT()/true] |
keep |
|
CometAggregateSuite |
1746 |
test final sum |
convert |
aggregate/group_by_numeric_keys.sql (new) |
CometAggregateSuite |
1767 |
regression aggregate flags follow the Spark patch release that changed them |
keep |
|
CometAggregateSuite |
1818 |
avg/sum overflow on decimal(38, _) |
convert |
aggregate/sum_avg_decimal.sql (new) |
CometAggregateSuite |
1862 |
decimal sum recovers from an intermediate overflow without group by |
convert |
aggregate/sum_decimal_max_precision.sql (new) |
CometAggregateSuite |
1885 |
decimal sum intermediate overflow with group by nulls like Spark |
split (rest: #6617) |
aggregate/sum_decimal_max_precision_legacy.sql (new) |
CometAggregateSuite |
1907 |
decimal sum partial overflow across partitions |
split (rest: #6617) |
aggregate/sum_decimal_max_precision_legacy.sql (new) |
CometAggregateSuite |
1934 |
decimal sum merged final overflow raises the out-of-range error |
split (rest: #6617) |
aggregate/sum_decimal_max_precision_legacy.sql (new) |
CometAggregateSuite |
1959 |
decimal sum without codegen falls back at maximum precision |
convert |
aggregate/sum_decimal_max_precision_legacy.sql (new) + per-config siblings |
CometAggregateSuite |
2023 |
decimal sum with a codegen fallback expression falls back at maximum precision |
convert |
aggregate/sum_decimal_max_precision_legacy.sql (new) + ANSI twin |
CometAggregateSuite |
2093 |
decimal sum with object hash aggregate falls back at maximum precision |
convert |
aggregate/sum_decimal_max_precision_legacy.sql (new) |
CometAggregateSuite |
2127 |
decimal sum with ungrouped object hash aggregate recovers like Spark |
convert |
aggregate/sum_decimal_max_precision.sql (new) |
CometAggregateSuite |
2152 |
decimal sum distinct recovers without group by and latches with group by |
split (rest: #6617) |
aggregate/sum_decimal_max_precision.sql (new) + sum_decimal_max_precision_legacy.sql (new) |
CometAggregateSuite |
2188 |
avg decimal with nothing to average |
convert |
aggregate/avg_ansi.sql (new) |
CometAggregateSuite |
2212 |
test final avg |
convert |
aggregate/group_by_floating_point.sql (new) |
CometAggregateSuite |
2234 |
final decimal avg |
convert |
aggregate/sum_avg_decimal.sql (new) |
CometAggregateSuite |
2268 |
AVG stays in Spark across a Spark shuffle |
keep |
|
CometAggregateSuite |
2281 |
avg null handling |
convert |
aggregate/group_by_floating_point.sql (new) |
CometAggregateSuite |
2301 |
Decimal Avg with DF |
split (rest: plan shape stays in Scala) |
aggregate/sum_avg_decimal.sql (new) |
CometAggregateSuite |
2346 |
distinct |
convert |
aggregate/distinct.sql (new) |
CometAggregateSuite |
2388 |
skip partial aggregation preserves post-shuffle distinct |
convert |
aggregate/skip_partial_aggregation.sql (new) |
CometAggregateSuite |
2422 |
skip partial aggregation is disabled by default |
keep |
|
CometAggregateSuite |
2449 |
skip partial aggregation admits only supported native shuffle plans |
keep |
|
CometAggregateSuite |
2525 |
first/last |
convert |
aggregate/first_last.sql |
CometAggregateSuite |
2563 |
test bool_and/bool_or |
convert |
aggregate/bool_and.sql (new) |
CometAggregateSuite |
2588 |
bitwise aggregate |
convert |
aggregate/bit_agg.sql |
CometAggregateSuite |
2665 |
covariance & correlation |
convert |
aggregate/covariance.sql; aggregate/corr.sql |
CometAggregateSuite |
2737 |
corr - nan/null |
convert |
aggregate/corr.sql |
CometAggregateSuite |
2762 |
var_pop and var_samp |
convert |
aggregate/variance.sql |
CometAggregateSuite |
2801 |
stddev_pop and stddev_samp |
convert |
aggregate/stddev.sql |
CometAggregateSuite |
2843 |
AVG and try_avg - basic functionality |
convert |
aggregate/avg_ansi.sql (new; also holds try_avg) |
CometAggregateSuite |
2879 |
AVG and try_avg - special numbers |
convert |
aggregate/avg_ansi.sql (new; also holds try_avg) |
CometAggregateSuite |
2926 |
ANSI support for sum - null test |
convert |
aggregate/sum_ansi.sql (new) |
CometAggregateSuite |
2939 |
ANSI support for decimal sum - null test |
convert |
aggregate/sum_ansi.sql (new) |
CometAggregateSuite |
2955 |
ANSI support for try_sum - null test |
convert |
aggregate/try_sum.sql (new) |
CometAggregateSuite |
2968 |
ANSI support for try_sum decimal - null test |
convert |
aggregate/try_sum.sql (new) |
CometAggregateSuite |
2984 |
ANSI support for sum - null test (group by) |
convert |
aggregate/sum_ansi.sql (new) |
CometAggregateSuite |
3003 |
ANSI support for decimal sum - null test (group by) |
convert |
aggregate/sum_ansi.sql (new) |
CometAggregateSuite |
3022 |
ANSI support for try_sum - null test (group by) |
convert |
aggregate/try_sum.sql (new) |
CometAggregateSuite |
3040 |
ANSI support for try_sum decimal - null test (group by) |
convert |
aggregate/try_sum.sql (new) |
CometAggregateSuite |
3063 |
ANSI support - SUM function |
convert |
aggregate/sum_ansi_overflow.sql (new); aggregate/sum.sql; aggregate/sum_ansi.sql (new) |
CometAggregateSuite |
3118 |
ANSI support for decimal SUM function |
convert |
aggregate/sum_ansi_overflow.sql (new); aggregate/sum.sql |
CometAggregateSuite |
3139 |
ANSI support for SUM - GROUP BY |
convert |
aggregate/sum_ansi_overflow.sql (new); aggregate/sum.sql; aggregate/sum_ansi.sql (new) |
CometAggregateSuite |
3198 |
ANSI support for decimal SUM - GROUP BY |
convert |
aggregate/sum_ansi_overflow.sql (new); aggregate/sum.sql |
CometAggregateSuite |
3220 |
try_sum overflow - with GROUP BY |
convert |
aggregate/try_sum.sql (new) |
CometAggregateSuite |
3260 |
try_sum decimal overflow |
convert |
aggregate/try_sum.sql (new) |
CometAggregateSuite |
3267 |
try_sum decimal overflow - with GROUP BY |
convert |
aggregate/try_sum.sql (new) |
CometAggregateSuite |
3274 |
try_sum decimal partial overflow - with GROUP BY |
convert |
aggregate/try_sum.sql (new) |
CometAggregateSuite |
3286 |
SumDecimal and AvgDecimal nullable should always be true |
convert |
aggregate/sum_ansi.sql (new); aggregate/avg_ansi.sql (new) |
CometAggregateSuite |
3351 |
group by array of map falls back to Spark (issue #4123) |
convert |
aggregate/group_by_map_fallback.sql (new) |
CometAggregateSuite |
3379 |
HashAggregate with catalyst-pruned resultExpressions returns 0-col output |
convert |
aggregate/hash_aggregate_pruned_output.sql (new) |
CometExpressionSuite |
4044 |
hll_sketch_agg and hll_sketch_estimate (incompatible, opt-in) |
convert |
aggregate/hll_sketch.sql (new) |
CometExpressionSuite |
4069 |
hll_union_agg and hll_union (incompatible, opt-in) |
convert |
aggregate/hll_sketch.sql (new) |
CometExpressionSuite |
4097 |
hll_union_agg rejects different lgConfigK when not allowed |
blocked: #6617 |
aggregate/hll_union_agg_lgk.sql (new) |
CometExpressionSuite |
4132 |
hll_union with a NULL allowDifferentLgConfigK returns NULL |
blocked: #6618 |
aggregate/hll_sketch.sql (new) |
CometExpressionSuite |
4175 |
hll_sketch_agg over all-null input estimates to 0, not NULL |
convert |
aggregate/hll_sketch.sql (new) |
CometFuzzAggregateSuite |
26 |
count distinct - simple columns |
keep |
|
CometFuzzAggregateSuite |
40 |
count distinct - complex columns |
keep |
|
CometFuzzAggregateSuite |
50 |
count distinct group by multiple column - simple columns |
keep |
|
CometFuzzAggregateSuite |
64 |
count distinct group by multiple column - complex columns |
keep |
|
CometFuzzAggregateSuite |
76 |
count distinct multiple values and group by multiple column |
keep |
|
CometFuzzAggregateSuite |
86 |
count(*) group by single column |
keep |
|
CometFuzzAggregateSuite |
96 |
count(col) group by single column |
keep |
|
CometFuzzAggregateSuite |
107 |
count(col1, col2, ..) group by single column |
keep |
|
CometFuzzAggregateSuite |
118 |
min/max aggregate |
keep |
|
Part of #6615.
What / Why
This issue moves aggregate function coverage from
exec/CometAggregateSuiteto Comet SQL tests undersql-tests/expressions/aggregate/. It also covers the HLL sketch tests inCometExpressionSuite. Aggregate operator behavior that only Scala can observe stays where it is. Static analysis onmainat 3bc2faa found the following:CometAggregateSuiteCometFuzzAggregateSuiteCometExpressionSuiteNone of these tests is fully covered by an existing fixture today. The fixtures are thinner in three ways:
tolerance=1e-6for avg over ints, grouped var/stddev, and the corr NaN/NULL cases. Tolerance also skips NaN (test:query tolerance=in Comet SQL tests passes when either side is NaN #6616).try_sum,try_avg,bool_and/every;spark.sql.legacy.statisticalAggregate;SELECT DISTINCT;spark.comet.shuffle.mode=jvm.How
checkSparkAnswerAndNumOfAggregates(q, n)maps. The helper ischeckSparkAnswer(answer only) plus a count ofCometHashAggregateExec. Whennis 2 (or 4 forCOUNT(DISTINCT)) on a fully native plan, a plainqueryis equal or stronger. Counts of 1 or 0 underspark.comet.shuffle.enabled=falseor a disabled partial/final mode test mixed-engine behavior; those parts stay in Scala.Batch shape. Under
local[5], anINSERT ... VALUESof a few rows writes one file per row, so each partial aggregate sees one row. For ordered or multi-row input useINSERT INTO t SELECT /*+ COALESCE(1) */ ...orrange(0, n, 1, 1). The Scala ORDER BY views become subqueries; the Sort survives under first/last because those aggregates are order-sensitive.Work
The work splits into three PR-sized parts. Line numbers refer to
CometAggregateSuiteunless marked.Part 1: grouping keys, types, DISTINCT, batches
group_by_numeric_keys.sql(new)MIN(v)+2,SUM(v+1),k+2andsum(v + g3).group_by_floating_point.sql(new): L276 (group part), L1175, L1190, L2212, L2281.float('-0.0')/double('-0.0')and NULL.min_max_strict_floating_point.sql(new): L276, withspark.comet.exec.strictFloatingPoint=true. Twoexpect_fallbackqueries per min/max over float/double zeros, one per reason.min_max_datetime.sql(new): L1303,-- ConfigMatrix: spark.sql.parquet.outputTimestampType=TIMESTAMP_MICROS,TIMESTAMP_MILLIS.distinct.sql(new): L686, L699, L840, L2346.count(DISTINCT)over a literal.count(DISTINCT)with a NULL in the tuple.group_by_multi_batch.sql(new): L1320, L1434, L1506.-- ConfigMatrix: spark.comet.batchSize=128,1024,10001.first_last_multi_batch.sql(new): L1402, L1477, L1571.first/lastwith IGNORE NULLS over a sorted subquery, and 8192 groups.group_by_all_types.sql(new): L1539, L1588. Keys of all 14 primitive types, with batch-size, dictionary and outputTimestampType matrices.collect_grouped_batches.sql(new): L191, withspark.comet.batchSize=128.collect_list.sqlcoalesce, and the two-stage collect_list -> named_struct -> collect_list shape.collect_list/collect_setofCAST(NULL AS ARRAY<STRUCT<...>>).bool_and.sql(new): L2563, coveringbool_and/everyand min/max over boolean.bit_agg.sql: L2588.bit_or/bit_xorover tinyint and smallint, grouped by a smallint key, under jvm shuffle.first_last.sql: L2525, plus-- ConfigMatrix: spark.comet.shuffle.mode=native,jvm.partial_merge.sqlsum(string)withcount(DISTINCT string)over 3 files.queryforavg(int)+count(DISTINCT),sum(1)+sum(DISTINCT)andavg(DISTINCT)+sum.group_by_map_fallback.sql(new): L1075, L3351, withMinSparkVersion: 4.1andspark.comet.exec.localTableScan.enabled=true.hash_aggregate_pruned_output.sql(new): L3379.Part 2: sum/avg (decimal, ANSI, try) and statistical aggregates
sum_avg_decimal.sql(new): L309, L949, L966, L1209, L1223, L1242, L1263 (part), L1818, L2234, L2301 (part).sum_decimal_max_precision.sql(new): L1862, L2127, L2152, with an ANSI matrix.range(0, n, 1, 1), to keep the running-sum order.sum_decimal_max_precision_legacy.sql(new), plus ANSI twins and per-codegen-config siblings: the non-ANSI halves of L1885, L1907, L1934 and L2152, plus L1959, L2023 and L2093.maxFields=1, andNO_CODEGENwith a 3.5+/3.4 version split..WITH_SUGGESTION).sum_ansi.sql(new, ANSI matrix): L2926, L2939, L2984, L3003, the non-overflow parts of L3063 and L3139, and the sum part of L3286.sum_ansi_overflow.sql(new, ANSI on): theexpect_error(ARITHMETIC_OVERFLOW)halves of L3063, L3118, L3139 and L3198. The Scala check (both engines throw, both messages contain the class) is exactly whatexpect_errordoes.try_sum.sql(new, ANSI matrix): L2955, L2968, L3022, L3040, L3220, L3260, L3267, L3274.avg_ansi.sql(new, ANSI matrix; also holdstry_avg): L2188, L2843, L2879, and the avg part of L3286.sum.sql/*+ REPARTITION(2) */hint for.repartition(2).avg.sql(L1050, part):AVG(int)grouped by a string key, compared exactly.stddev.sql/variance.sql(L671, L2762, L2801)-- ConfigMatrix: spark.sql.legacy.statisticalAggregate=true,false, a shuffle-mode matrix and a dictionary matrix.float('Infinity')case.var_pop(col5). The Scala test meant to cover it but runsvar_samp(col5)instead.covariance.sql,corr.sql(L2665, L2737)queryfor NULL and NaN results; use tolerance with ORDER BY only for finite results.hll_sketch.sql(new,MinSparkVersion: 4.0):CometExpressionSuiteL4044, L4069 and L4175.abs(est - 700) / 700.0 <= 0.05. A plainquerythen asserts native execution and the 5% bound without comparing estimates bit for bit.Part 3: operator-level tests that happen to be expressible (optional)
These tests assert only answers or fallback reasons, and every config they need is a SQL conf, so they can move. They carry no expression coverage, so migrating them is optional.
mixed_engine_spark_partial.sql: L718, L731. AVALUESsource keeps the Partial in Spark.mixed_engine_partial_disabled.sql: L800.approx_percentile_distinct.sql: L340, with a matrix overspark.comet.testing.aggregate.partialMode.enabled/finalMode.enabled.skip_partial_aggregation.sql: L2388.correlated_subquery_aggregate.sql: L874, the TPC-H Q17 shape.collect_distinct_fallback.sql: L240.Blocked
CometExpressionSuiteL4097 (hll_union_agglgConfigK). Each asserts native execution plus error-class, SQLSTATE and exception-class parity.EliminateSortsexcluded, so the Sort below the SUM survives.CometExpressionSuiteL4132 needsNullPropagationexcluded.Stays in Scala
ignored with "TODO: To be addressed after DF 55 migration" and no link. Re-enable the two ignored CometAggregateSuite metric tests after the DataFusion 55 peak_mem_used change #5703 trackspeak_mem_used, so link it from the ignore.CometFilterExecpresence and materialized AQE stages.java_methodplan check in L1657, stay for the same reason.CometFuzzAggregateSuite(all 9). It is the fuzz layer: it sweeps whatever schema the generator produces and picks up new types automatically. The fixtures should still close the gaps it exposes. No fixture asserts nativecount(DISTINCT)for boolean, tinyint, smallint, float, decimal(36,18), date, timestamp, timestamp_ntz or binary.Problems found along the way
Vacuous or misleading Scala tests
null.asInstanceOf[Int], which is0in Scala. Their expected rows are reallyRow(0, ...), and "with all nulls" aggregates zeros.CometSum.supportsNativePartialToSparkFinalis false for decimal, so the Partial is restored to Spark.spark.comet.exec.shuffle.fallbackToColumnar, which is not defined anywhere inspark/src/main.var_popover double.df.schemacomes from Spark's analyzed plan.EliminateSortsremoves it below an order-insensitive aggregate.Fixture problems
aggregate_filter.sqllinks to the placeholderissues/XXXX, and lines 44-45 and 60-61 are the same query.min_max.sql's only ungrouped query mixes in string min/max, so it isexpect_fallback(SortAggregate is not supported). Native ungrouped min/max over int or double is never asserted.first_last.sql("large group (multi-batch)") andpartial_merge.sql("large, multi-batch input") write 1000 rows as one file, below the 8192 default batch size, so both run as a single batch.corr.sqltest_corr_nan,avg.sqltest_avg_specialandsum.sql's Infinity rows compare these results only under tolerance, which skips them (test:query tolerance=in Comet SQL tests passes when either side is NaN #6616).collect_list.sqlhas groups mixing0.0and-0.0undersort_array, which is order-sensitive.partial_merge.sqland L2301 setspark.comet.expression.Cast.allowIncompatible=true, butCometCast.canCastFromDoublenow reports DecimalType as Compatible. This was not verified by running.Done when
.github/workflows/pr_build_linux.ymlandpr_build_macos.yml.run-all-spark-profiles, because pull requests run only the default profile.Every Scala test in scope, with its verdict and target fixture
CometAggregateSuiteCometAggregateSuiteCometAggregateSuiteCometAggregateSuiteCometAggregateSuiteCometAggregateSuiteCometAggregateSuiteCometAggregateSuiteCometAggregateSuiteCometAggregateSuiteCometAggregateSuiteCometAggregateSuiteCometAggregateSuiteCometAggregateSuiteCometAggregateSuiteCometAggregateSuiteCometAggregateSuiteCometAggregateSuiteCometAggregateSuiteCometAggregateSuiteCometAggregateSuiteCometAggregateSuiteCometAggregateSuiteCometAggregateSuiteCometAggregateSuiteCometAggregateSuiteCometAggregateSuiteCometAggregateSuiteCometAggregateSuiteCometAggregateSuiteCometAggregateSuiteCometAggregateSuiteCometAggregateSuiteCometAggregateSuiteCometAggregateSuiteCometAggregateSuiteCometAggregateSuiteCometAggregateSuiteCometAggregateSuiteCometAggregateSuiteCometAggregateSuiteCometAggregateSuiteCometAggregateSuiteCometAggregateSuiteCometAggregateSuiteCometAggregateSuiteCometAggregateSuiteCometAggregateSuiteCometAggregateSuiteCometAggregateSuiteCometAggregateSuiteCometAggregateSuiteCometAggregateSuiteCometAggregateSuiteCometAggregateSuiteCometAggregateSuiteCometAggregateSuiteCometAggregateSuiteCometAggregateSuiteCometAggregateSuiteCometAggregateSuiteCometAggregateSuiteCometAggregateSuiteCometAggregateSuiteCometAggregateSuiteCometAggregateSuiteCometAggregateSuiteCometAggregateSuiteCometAggregateSuiteCometAggregateSuiteCometAggregateSuiteCometAggregateSuiteCometAggregateSuiteCometAggregateSuiteCometAggregateSuiteCometAggregateSuiteCometAggregateSuiteCometAggregateSuiteCometAggregateSuiteCometAggregateSuiteCometAggregateSuiteCometAggregateSuiteCometAggregateSuiteCometAggregateSuiteCometAggregateSuiteCometAggregateSuiteCometAggregateSuiteCometAggregateSuiteCometAggregateSuiteCometAggregateSuiteCometAggregateSuiteCometAggregateSuiteCometAggregateSuiteCometAggregateSuiteCometAggregateSuiteCometAggregateSuiteCometAggregateSuiteCometAggregateSuiteCometAggregateSuiteCometAggregateSuiteCometAggregateSuiteCometAggregateSuiteCometAggregateSuiteCometAggregateSuiteCometAggregateSuiteCometAggregateSuiteCometAggregateSuiteCometAggregateSuiteCometAggregateSuiteCometAggregateSuiteCometAggregateSuiteCometAggregateSuiteCometAggregateSuiteCometExpressionSuiteCometExpressionSuiteCometExpressionSuiteCometExpressionSuiteCometExpressionSuiteCometFuzzAggregateSuiteCometFuzzAggregateSuiteCometFuzzAggregateSuiteCometFuzzAggregateSuiteCometFuzzAggregateSuiteCometFuzzAggregateSuiteCometFuzzAggregateSuiteCometFuzzAggregateSuiteCometFuzzAggregateSuite