Skip to content

test: move aggregate function coverage from CometAggregateSuite to Comet SQL tests #6631

Description

@andygrove

Part of #6615.

What / Why

This issue moves aggregate function coverage from exec/CometAggregateSuite to Comet SQL tests under sql-tests/expressions/aggregate/. It also covers the HLL sketch tests in CometExpressionSuite. Aggregate operator behavior that only Scala can observe stays where it is. Static analysis on main at 3bc2faa found the following:

Suite Covered Convert Blocked Split Keep Total
CometAggregateSuite 0 87 2 8 16 113
CometFuzzAggregateSuite 0 0 0 0 9 9
CometExpressionSuite 0 3 2 0 0 5
Total 0 90 4 8 25 127

None of these tests is fully covered by an existing fixture today. The fixtures are thinner in three ways:

  1. String group keys only. Every grouped fixture for sum, count, min/max, avg, bit_agg, stddev, variance, covariance, corr, first_last and collect_* groups by a string key. The Scala tests group by tinyint, smallint, int, bigint, float, double, decimal, date, timestamp, binary and VARCHAR keys. DataFusion uses a different group-values implementation for a single primitive key than for a single string key.
  2. Tolerance where Scala compares exactly. The fixtures use tolerance=1e-6 for avg over ints, grouped var/stddev, and the corr NaN/NULL cases. Tolerance also skips NaN (test: query tolerance= in Comet SQL tests passes when either side is NaN #6616).
  3. Whole areas have no fixture:
    • try_sum, try_avg, bool_and / every;
    • sum/avg under ANSI, and decimal SUM overflow at maximum precision;
    • spark.sql.legacy.statisticalAggregate;
    • strictFloatingPoint min/max;
    • GROUP BY without aggregate functions, and SELECT DISTINCT;
    • min/max over float, decimal, date, timestamp or boolean;
    • aggregates under spark.comet.shuffle.mode=jvm.

How checkSparkAnswerAndNumOfAggregates(q, n) maps. The helper is checkSparkAnswer (answer only) plus a count of CometHashAggregateExec. When n is 2 (or 4 for COUNT(DISTINCT)) on a fully native plan, a plain query is equal or stronger. Counts of 1 or 0 under spark.comet.shuffle.enabled=false or a disabled partial/final mode test mixed-engine behavior; those parts stay in Scala.

Batch shape. Under local[5], an INSERT ... VALUES of a few rows writes one file per row, so each partial aggregate sees one row. For ordered or multi-row input use INSERT INTO t SELECT /*+ COALESCE(1) */ ... or range(0, n, 1, 1). The Scala ORDER BY views become subqueries; the Sort survives under first/last because those aggregates are order-sensitive.

Work

The work splits into three PR-sized parts. Line numbers refer to CometAggregateSuite unless marked.

Part 1: grouping keys, types, DISTINCT, batches

  • group_by_numeric_keys.sql (new)
    • Covers L856, L982, L999, L1012, L1037, L1102, L1123, L1147, L1619, L1636, L1657 (the convertible part) and L1746.
    • Use int and bigint keys with real NULLs; see the note on L1102-L1147 below.
    • SUM/COUNT/MIN/MAX/AVG grouped and ungrouped, with exact avg.
    • Result expressions such as MIN(v)+2, SUM(v+1), k+2 and sum(v + g3).
    • Add shuffle-mode and dictionary ConfigMatrix lines.
  • group_by_floating_point.sql (new): L276 (group part), L1175, L1190, L2212, L2281.
    • Float and double keys with NaN, float('-0.0') / double('-0.0') and NULL.
    • Mostly singleton groups, so SUM/AVG compare exactly.
  • min_max_strict_floating_point.sql (new): L276, with spark.comet.exec.strictFloatingPoint=true. Two expect_fallback queries per min/max over float/double zeros, one per reason.
  • min_max_datetime.sql (new): L1303, -- ConfigMatrix: spark.sql.parquet.outputTimestampType=TIMESTAMP_MICROS,TIMESTAMP_MILLIS.
  • distinct.sql (new): L686, L699, L840, L2346.
    • Multi-group count(DISTINCT) over a literal.
    • Multi-column count(DISTINCT) with a NULL in the tuple.
    • GROUP BY without aggregate functions; SELECT DISTINCT; repeated distinct aggregates.
    • Shuffle-mode ConfigMatrix.
  • group_by_multi_batch.sql (new): L1320, L1434, L1506.
    • -- ConfigMatrix: spark.comet.batchSize=128,1024,10001.
    • Deterministic keys at cardinality 1, 100 and 10000, including negatives.
  • first_last_multi_batch.sql (new): L1402, L1477, L1571. first/last with IGNORE NULLS over a sorted subquery, and 8192 groups.
  • group_by_all_types.sql (new): L1539, L1588. Keys of all 14 primitive types, with batch-size, dictionary and outputTimestampType matrices.
  • collect_grouped_batches.sql (new): L191, with spark.comet.batchSize=128.
  • collect_list.sql
    • L157: a struct with a non-nullable field via coalesce, and the two-stage collect_list -> named_struct -> collect_list shape.
    • L215: collect_list / collect_set of CAST(NULL AS ARRAY<STRUCT<...>>).
  • bool_and.sql (new): L2563, covering bool_and / every and min/max over boolean.
  • bit_agg.sql: L2588. bit_or / bit_xor over tinyint and smallint, grouped by a smallint key, under jvm shuffle.
  • first_last.sql: L2525, plus -- ConfigMatrix: spark.comet.shuffle.mode=native,jvm.
  • partial_merge.sql
    • L1347: sum(string) with count(DISTINCT string) over 3 files.
    • L1388: exact query for avg(int)+count(DISTINCT), sum(1)+sum(DISTINCT) and avg(DISTINCT)+sum.
  • group_by_map_fallback.sql (new): L1075, L3351, with MinSparkVersion: 4.1 and spark.comet.exec.localTableScan.enabled=true.
  • hash_aggregate_pruned_output.sql (new): L3379.

Part 2: sum/avg (decimal, ANSI, try) and statistical aggregates

  • sum_avg_decimal.sql (new): L309, L949, L966, L1209, L1223, L1242, L1263 (part), L1818, L2234, L2301 (part).
    • Columns: decimal(5,2), decimal(18,10) and decimal(38,37) with NULL rows; tinyint/smallint/int/bigint keys.
    • MIN/MAX/SUM/AVG, including decimal(38,37) overflow to NULL.
  • sum_decimal_max_precision.sql (new): L1862, L2127, L2152, with an ANSI matrix.
    • decimal(38,38) rows 0.6, 0.6, -0.6 written as one ordered file via range(0, n, 1, 1), to keep the running-sum order.
  • sum_decimal_max_precision_legacy.sql (new), plus ANSI twins and per-codegen-config siblings: the non-ANSI halves of L1885, L1907, L1934 and L2152, plus L1959, L2023 and L2093.
    • L1959 fans out to about 8 files: whole-stage codegen off, maxFields=1, and NO_CODEGEN with a 3.5+/3.4 version split.
    • L1934's error class changes at 4.0 (.WITH_SUGGESTION).
  • sum_ansi.sql (new, ANSI matrix): L2926, L2939, L2984, L3003, the non-overflow parts of L3063 and L3139, and the sum part of L3286.
  • sum_ansi_overflow.sql (new, ANSI on): the expect_error(ARITHMETIC_OVERFLOW) halves of L3063, L3118, L3139 and L3198. The Scala check (both engines throw, both messages contain the class) is exactly what expect_error does.
  • try_sum.sql (new, ANSI matrix): L2955, L2968, L3022, L3040, L3220, L3260, L3267, L3274.
  • avg_ansi.sql (new, ANSI matrix; also holds try_avg): L2188, L2843, L2879, and the avg part of L3286.
  • sum.sql
    • L1028: grouped bigint overflow.
    • L1467: a VARCHAR(3) group key.
    • The non-ANSI halves of L3063, L3118, L3139 and L3198. Use a /*+ REPARTITION(2) */ hint for .repartition(2).
  • avg.sql (L1050, part): AVG(int) grouped by a string key, compared exactly.
  • stddev.sql / variance.sql (L671, L2762, L2801)
    • Add -- ConfigMatrix: spark.sql.legacy.statisticalAggregate=true,false, a shuffle-mode matrix and a dictionary matrix.
    • Add a one-row float('Infinity') case.
    • Add var_pop(col5). The Scala test meant to cover it but runs var_samp(col5) instead.
  • covariance.sql, corr.sql (L2665, L2737)
    • The nine datasets, with legacy, shuffle and dictionary matrices.
    • Use exact query for NULL and NaN results; use tolerance with ORDER BY only for finite results.
  • hll_sketch.sql (new, MinSparkVersion: 4.0): CometExpressionSuite L4044, L4069 and L4175.
    • Wrap each estimate in the test's own bound, e.g. abs(est - 700) / 700.0 <= 0.05. A plain query then asserts native execution and the 5% bound without comparing estimates bit for bit.

Part 3: operator-level tests that happen to be expressible (optional)

These tests assert only answers or fallback reasons, and every config they need is a SQL conf, so they can move. They carry no expression coverage, so migrating them is optional.

  • mixed_engine_spark_partial.sql: L718, L731. A VALUES source keeps the Partial in Spark.
  • mixed_engine_partial_disabled.sql: L800.
  • approx_percentile_distinct.sql: L340, with a matrix over spark.comet.testing.aggregate.partialMode.enabled / finalMode.enabled.
  • skip_partial_aggregation.sql: L2388.
  • correlated_subquery_aggregate.sql: L874, the TPC-H Q17 shape.
  • collect_distinct_fallback.sql: L240.

Blocked

Stays in Scala

  • Metrics: L63, L105, L2422, L2449.
  • Mixed-engine buffer safety and plan shape: L359, L388, L455, L512, L540, L580, L746, L786, L823, L2268. These check unsafe-partial tags, Comet aggregate counts of 0 or 1, CometFilterExec presence and materialized AQE stages.
    • The shuffle-off halves of L1050, L1263 and L2301, and the java_method plan check in L1657, stay for the same reason.
  • Plan canonicalization and exchange reuse: L1698.
  • A helper unit test with no query: L1767.
  • CometFuzzAggregateSuite (all 9). It is the fuzz layer: it sweeps whatever schema the generator produces and picks up new types automatically. The fixtures should still close the gaps it exposes. No fixture asserts native count(DISTINCT) for boolean, tinyint, smallint, float, decimal(36,18), date, timestamp, timestamp_ntz or binary.

Problems found along the way

Vacuous or misleading Scala tests

  • The "null" key tests have no NULLs. L1102, L1123 and L1147 build keys with null.asInstanceOf[Int], which is 0 in Scala. Their expected rows are really Row(0, ...), and "with all nulls" aggregates zeros.
  • L949 no longer runs Comet's decimal SUM. With Comet shuffle off the Final is in Spark, and CometSum.supportsNativePartialToSparkFinal is false for decimal, so the Partial is restored to Spark.
  • L1347 sets spark.comet.exec.shuffle.fallbackToColumnar, which is not defined anywhere in spark/src/main.
  • L686, "count with aggregation filter", has no FILTER clause.
  • L2762 never tests var_pop over double.
  • L3286's nullability assertion is vacuous, because df.schema comes from Spark's analyzed plan.
  • L1320's view ORDER BY is dropped. EliminateSorts removes it below an order-insensitive aggregate.

Fixture problems

  • aggregate_filter.sql links to the placeholder issues/XXXX, and lines 44-45 and 60-61 are the same query.
  • min_max.sql's only ungrouped query mixes in string min/max, so it is expect_fallback(SortAggregate is not supported). Native ungrouped min/max over int or double is never asserted.
  • "Multi-batch" comments are wrong. first_last.sql ("large group (multi-batch)") and partial_merge.sql ("large, multi-batch input") write 1000 rows as one file, below the 8192 default batch size, so both run as a single batch.
  • NaN and infinity under tolerance. corr.sql test_corr_nan, avg.sql test_avg_special and sum.sql's Infinity rows compare these results only under tolerance, which skips them (test: query tolerance= in Comet SQL tests passes when either side is NaN #6616).
  • collect_list.sql has groups mixing 0.0 and -0.0 under sort_array, which is order-sensitive.
  • Possibly stale config. partial_merge.sql and L2301 set spark.comet.expression.Cast.allowIncompatible=true, but CometCast.canCastFromDouble now reports DecimalType as Compatible. This was not verified by running.

Done when

  • Every test in the table below is accounted for. Either a fixture covers it (the same inputs, configs and assertion strength, or stronger), the fixture that already covers it is cited, or it stays in Scala for the reason given above.
  • Each converted Scala test is deleted in the PR that adds its coverage. A suite left empty is deleted and removed from the suite lists in .github/workflows/pr_build_linux.yml and pr_build_macos.yml.
  • The new and changed fixtures pass on every Spark profile. Apply run-all-spark-profiles, because pull requests run only the default profile.
Every Scala test in scope, with its verdict and target fixture
Suite Line Test Verdict Target
CometAggregateSuite 63 grouped aggregate metrics are forwarded without fabricating global metrics (ignored) keep
CometAggregateSuite 105 range sampling does not report grouped aggregate metrics (ignored) keep
CometAggregateSuite 157 collect_list over struct with non-nullable fields convert aggregate/collect_list.sql
CometAggregateSuite 191 grouped collect_list/collect_set over nulls, duplicates and several batches convert aggregate/collect_grouped_batches.sql (new)
CometAggregateSuite 215 collect_list/collect_set over typed nested NULL (grouped=$grouped) [instances: grouped=false, grouped=true] convert aggregate/collect_list.sql
CometAggregateSuite 240 collect_list/collect_set combined with distinct aggregate falls back safely convert aggregate/collect_distinct_fallback.sql (new)
CometAggregateSuite 276 min/max floating point with negative zero convert aggregate/min_max_strict_floating_point.sql (new); aggregate/group_by_floating_point.sql (new)
CometAggregateSuite 309 avg decimal convert aggregate/sum_avg_decimal.sql (new)
CometAggregateSuite 340 approx_percentile with distinct aggregate does not split across Comet and Spark convert aggregate/approx_percentile_distinct.sql (new)
CometAggregateSuite 359 disabled FIRST/LAST preserves percentile buffers with local Comet shuffle keep
CometAggregateSuite 388 decimal AVG falls back across a Spark shuffle (AQE=$adaptive) [instances: AQE=false, AQE=true] keep
CometAggregateSuite 455 COUNT and AVG fall back together across a Spark shuffle (AQE=$adaptive) [instances: AQE=false, AQE=true] keep
CometAggregateSuite 512 COUNT preserves safe native partials across a Spark shuffle (AQE=$adaptive) [instances: AQE=false, AQE=true] keep
CometAggregateSuite 540 $fn falls back when enabled native shuffle is ineligible (AQE=$adaptive) [instances: collect_list/collect_set x AQE=false/true] keep
CometAggregateSuite 580 $fn preserves aggregate buffers with an unsupported array hash key (AQE=$adaptive) [instances: percentile/collect_list/sum x AQE=false/true] keep
CometAggregateSuite 671 stddev_pop should return NaN for some cases convert aggregate/stddev.sql
CometAggregateSuite 686 count with aggregation filter convert aggregate/distinct.sql (new)
CometAggregateSuite 699 multiple column distinct count convert aggregate/distinct.sql (new)
CometAggregateSuite 718 Only trigger Comet Final aggregation on Comet partial aggregation convert aggregate/mixed_engine_spark_partial.sql (new)
CometAggregateSuite 731 Average expression in Comet Final should handle all null inputs from partial Spark aggregation convert aggregate/mixed_engine_spark_partial.sql (new)
CometAggregateSuite 746 decimal SUM partial stays in Spark when a later input cancels precision overflow keep
CometAggregateSuite 786 mixed engine sum/avg falls back when Spark Final would consume native AVG keep
CometAggregateSuite 800 mixed engine sum/avg: Spark partial + Comet final matches Spark convert aggregate/mixed_engine_partial_disabled.sql (new)
CometAggregateSuite 823 mixed engine collect_list: $name matches Spark [instances: 'Comet partial + Spark final', 'Spark partial + Comet final'] keep
CometAggregateSuite 840 Aggregation without aggregate expressions should use correct result expressions convert aggregate/distinct.sql (new)
CometAggregateSuite 856 Final aggregation should not bind to the input of partial aggregation convert aggregate/group_by_numeric_keys.sql (new)
CometAggregateSuite 874 Ensure traversed operators during finding first partial aggregation are all native convert aggregate/correlated_subquery_aggregate.sql (new)
CometAggregateSuite 911 SUM decimal supports emit.first blocked: #6618 aggregate/sum_avg_decimal_sorted_input.sql (new)
CometAggregateSuite 930 AVG decimal supports emit.first blocked: #6618 aggregate/sum_avg_decimal_sorted_input.sql (new)
CometAggregateSuite 949 Fix NPE in partial decimal sum convert aggregate/sum_avg_decimal.sql (new)
CometAggregateSuite 966 fix: Decimal Average should not enable native final aggregation convert aggregate/sum_avg_decimal.sql (new)
CometAggregateSuite 982 trivial case convert aggregate/group_by_numeric_keys.sql (new)
CometAggregateSuite 999 avg convert aggregate/group_by_numeric_keys.sql (new)
CometAggregateSuite 1012 count, avg with null convert aggregate/group_by_numeric_keys.sql (new)
CometAggregateSuite 1028 SUM/AVG non-decimal overflow convert aggregate/sum.sql
CometAggregateSuite 1037 simple SUM, COUNT, MIN, MAX, AVG with non-distinct group keys convert aggregate/group_by_numeric_keys.sql (new)
CometAggregateSuite 1050 group-by on variable length types split (rest: plan shape stays in Scala) aggregate/avg.sql
CometAggregateSuite 1075 grouping on struct containing map should fallback to Spark convert aggregate/group_by_map_fallback.sql (new)
CometAggregateSuite 1102 simple SUM, COUNT, MIN, MAX, AVG with non-distinct + null group keys convert aggregate/group_by_numeric_keys.sql (new)
CometAggregateSuite 1123 simple SUM, COUNT, MIN, MAX, AVG with null aggregates convert aggregate/group_by_numeric_keys.sql (new)
CometAggregateSuite 1147 simple SUM, MIN, MAX, AVG with all nulls convert aggregate/group_by_numeric_keys.sql (new)
CometAggregateSuite 1175 SUM, COUNT, MIN, MAX, AVG on float & double convert aggregate/group_by_floating_point.sql (new)
CometAggregateSuite 1190 SUM, MIN, MAX, AVG for NaN, -0.0 and 0.0 convert aggregate/group_by_floating_point.sql (new)
CometAggregateSuite 1209 SUM/MIN/MAX/AVG on decimal convert aggregate/sum_avg_decimal.sql (new)
CometAggregateSuite 1223 multiple SUM/MIN/MAX/AVG on decimal and non-decimal convert aggregate/sum_avg_decimal.sql (new)
CometAggregateSuite 1242 SUM/AVG on decimal with different precisions convert aggregate/sum_avg_decimal.sql (new)
CometAggregateSuite 1263 SUM decimal with DF split (rest: plan shape stays in Scala) aggregate/sum_avg_decimal.sql (new)
CometAggregateSuite 1303 COUNT/MIN/MAX on date, timestamp convert aggregate/min_max_datetime.sql (new)
CometAggregateSuite 1320 single group-by column + aggregate column, multiple batches, no null convert aggregate/group_by_multi_batch.sql (new)
CometAggregateSuite 1347 partialMerge - cnt distinct + sum convert aggregate/partial_merge.sql
CometAggregateSuite 1388 partialMerge - distinct + non-distinct aggregates (Expand pattern) convert aggregate/partial_merge.sql
CometAggregateSuite 1402 multiple group-by columns + single aggregate column (first/last), with nulls convert aggregate/first_last_multi_batch.sql (new)
CometAggregateSuite 1434 multiple group-by columns + single aggregate column, with nulls convert aggregate/group_by_multi_batch.sql (new)
CometAggregateSuite 1467 string should be supported convert aggregate/sum.sql
CometAggregateSuite 1477 multiple group-by columns + multiple aggregate column (first/last), with nulls convert aggregate/first_last_multi_batch.sql (new)
CometAggregateSuite 1506 multiple group-by columns + multiple aggregate column, with nulls convert aggregate/group_by_multi_batch.sql (new)
CometAggregateSuite 1539 all types first/last, with nulls convert aggregate/group_by_all_types.sql (new)
CometAggregateSuite 1571 first/last with ignore null convert aggregate/first_last_multi_batch.sql (new)
CometAggregateSuite 1588 all types, with nulls convert aggregate/group_by_all_types.sql (new)
CometAggregateSuite 1619 test final count convert aggregate/group_by_numeric_keys.sql (new)
CometAggregateSuite 1636 test final min/max convert aggregate/group_by_numeric_keys.sql (new)
CometAggregateSuite 1657 test final min/max/count with result expressions split (rest: plan shape stays in Scala) aggregate/group_by_numeric_keys.sql (new)
CometAggregateSuite 1698 aggregate canonicalization preserves result expressions and equivalent reuse: $function, AQE=$adaptive [instances: COUNT()/false, COUNT(DISTINCT _2)/false, COUNT(DISTINCT _2) + SUM(_2)/false, CAST(SIZE(COLLECT_SET(_2)) AS BIGINT)/false, COUNT()/true] keep
CometAggregateSuite 1746 test final sum convert aggregate/group_by_numeric_keys.sql (new)
CometAggregateSuite 1767 regression aggregate flags follow the Spark patch release that changed them keep
CometAggregateSuite 1818 avg/sum overflow on decimal(38, _) convert aggregate/sum_avg_decimal.sql (new)
CometAggregateSuite 1862 decimal sum recovers from an intermediate overflow without group by convert aggregate/sum_decimal_max_precision.sql (new)
CometAggregateSuite 1885 decimal sum intermediate overflow with group by nulls like Spark split (rest: #6617) aggregate/sum_decimal_max_precision_legacy.sql (new)
CometAggregateSuite 1907 decimal sum partial overflow across partitions split (rest: #6617) aggregate/sum_decimal_max_precision_legacy.sql (new)
CometAggregateSuite 1934 decimal sum merged final overflow raises the out-of-range error split (rest: #6617) aggregate/sum_decimal_max_precision_legacy.sql (new)
CometAggregateSuite 1959 decimal sum without codegen falls back at maximum precision convert aggregate/sum_decimal_max_precision_legacy.sql (new) + per-config siblings
CometAggregateSuite 2023 decimal sum with a codegen fallback expression falls back at maximum precision convert aggregate/sum_decimal_max_precision_legacy.sql (new) + ANSI twin
CometAggregateSuite 2093 decimal sum with object hash aggregate falls back at maximum precision convert aggregate/sum_decimal_max_precision_legacy.sql (new)
CometAggregateSuite 2127 decimal sum with ungrouped object hash aggregate recovers like Spark convert aggregate/sum_decimal_max_precision.sql (new)
CometAggregateSuite 2152 decimal sum distinct recovers without group by and latches with group by split (rest: #6617) aggregate/sum_decimal_max_precision.sql (new) + sum_decimal_max_precision_legacy.sql (new)
CometAggregateSuite 2188 avg decimal with nothing to average convert aggregate/avg_ansi.sql (new)
CometAggregateSuite 2212 test final avg convert aggregate/group_by_floating_point.sql (new)
CometAggregateSuite 2234 final decimal avg convert aggregate/sum_avg_decimal.sql (new)
CometAggregateSuite 2268 AVG stays in Spark across a Spark shuffle keep
CometAggregateSuite 2281 avg null handling convert aggregate/group_by_floating_point.sql (new)
CometAggregateSuite 2301 Decimal Avg with DF split (rest: plan shape stays in Scala) aggregate/sum_avg_decimal.sql (new)
CometAggregateSuite 2346 distinct convert aggregate/distinct.sql (new)
CometAggregateSuite 2388 skip partial aggregation preserves post-shuffle distinct convert aggregate/skip_partial_aggregation.sql (new)
CometAggregateSuite 2422 skip partial aggregation is disabled by default keep
CometAggregateSuite 2449 skip partial aggregation admits only supported native shuffle plans keep
CometAggregateSuite 2525 first/last convert aggregate/first_last.sql
CometAggregateSuite 2563 test bool_and/bool_or convert aggregate/bool_and.sql (new)
CometAggregateSuite 2588 bitwise aggregate convert aggregate/bit_agg.sql
CometAggregateSuite 2665 covariance & correlation convert aggregate/covariance.sql; aggregate/corr.sql
CometAggregateSuite 2737 corr - nan/null convert aggregate/corr.sql
CometAggregateSuite 2762 var_pop and var_samp convert aggregate/variance.sql
CometAggregateSuite 2801 stddev_pop and stddev_samp convert aggregate/stddev.sql
CometAggregateSuite 2843 AVG and try_avg - basic functionality convert aggregate/avg_ansi.sql (new; also holds try_avg)
CometAggregateSuite 2879 AVG and try_avg - special numbers convert aggregate/avg_ansi.sql (new; also holds try_avg)
CometAggregateSuite 2926 ANSI support for sum - null test convert aggregate/sum_ansi.sql (new)
CometAggregateSuite 2939 ANSI support for decimal sum - null test convert aggregate/sum_ansi.sql (new)
CometAggregateSuite 2955 ANSI support for try_sum - null test convert aggregate/try_sum.sql (new)
CometAggregateSuite 2968 ANSI support for try_sum decimal - null test convert aggregate/try_sum.sql (new)
CometAggregateSuite 2984 ANSI support for sum - null test (group by) convert aggregate/sum_ansi.sql (new)
CometAggregateSuite 3003 ANSI support for decimal sum - null test (group by) convert aggregate/sum_ansi.sql (new)
CometAggregateSuite 3022 ANSI support for try_sum - null test (group by) convert aggregate/try_sum.sql (new)
CometAggregateSuite 3040 ANSI support for try_sum decimal - null test (group by) convert aggregate/try_sum.sql (new)
CometAggregateSuite 3063 ANSI support - SUM function convert aggregate/sum_ansi_overflow.sql (new); aggregate/sum.sql; aggregate/sum_ansi.sql (new)
CometAggregateSuite 3118 ANSI support for decimal SUM function convert aggregate/sum_ansi_overflow.sql (new); aggregate/sum.sql
CometAggregateSuite 3139 ANSI support for SUM - GROUP BY convert aggregate/sum_ansi_overflow.sql (new); aggregate/sum.sql; aggregate/sum_ansi.sql (new)
CometAggregateSuite 3198 ANSI support for decimal SUM - GROUP BY convert aggregate/sum_ansi_overflow.sql (new); aggregate/sum.sql
CometAggregateSuite 3220 try_sum overflow - with GROUP BY convert aggregate/try_sum.sql (new)
CometAggregateSuite 3260 try_sum decimal overflow convert aggregate/try_sum.sql (new)
CometAggregateSuite 3267 try_sum decimal overflow - with GROUP BY convert aggregate/try_sum.sql (new)
CometAggregateSuite 3274 try_sum decimal partial overflow - with GROUP BY convert aggregate/try_sum.sql (new)
CometAggregateSuite 3286 SumDecimal and AvgDecimal nullable should always be true convert aggregate/sum_ansi.sql (new); aggregate/avg_ansi.sql (new)
CometAggregateSuite 3351 group by array of map falls back to Spark (issue #4123) convert aggregate/group_by_map_fallback.sql (new)
CometAggregateSuite 3379 HashAggregate with catalyst-pruned resultExpressions returns 0-col output convert aggregate/hash_aggregate_pruned_output.sql (new)
CometExpressionSuite 4044 hll_sketch_agg and hll_sketch_estimate (incompatible, opt-in) convert aggregate/hll_sketch.sql (new)
CometExpressionSuite 4069 hll_union_agg and hll_union (incompatible, opt-in) convert aggregate/hll_sketch.sql (new)
CometExpressionSuite 4097 hll_union_agg rejects different lgConfigK when not allowed blocked: #6617 aggregate/hll_union_agg_lgk.sql (new)
CometExpressionSuite 4132 hll_union with a NULL allowDifferentLgConfigK returns NULL blocked: #6618 aggregate/hll_sketch.sql (new)
CometExpressionSuite 4175 hll_sketch_agg over all-null input estimates to 0, not NULL convert aggregate/hll_sketch.sql (new)
CometFuzzAggregateSuite 26 count distinct - simple columns keep
CometFuzzAggregateSuite 40 count distinct - complex columns keep
CometFuzzAggregateSuite 50 count distinct group by multiple column - simple columns keep
CometFuzzAggregateSuite 64 count distinct group by multiple column - complex columns keep
CometFuzzAggregateSuite 76 count distinct multiple values and group by multiple column keep
CometFuzzAggregateSuite 86 count(*) group by single column keep
CometFuzzAggregateSuite 96 count(col) group by single column keep
CometFuzzAggregateSuite 107 count(col1, col2, ..) group by single column keep
CometFuzzAggregateSuite 118 min/max aggregate keep
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    area:aggregationHash aggregates, aggregate expressionsenhancementNew feature or requesttestTesting related

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions