You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
{{ message }}
Repository navigation
[EPIC] Move expression test coverage from Scala suites to Comet SQL tests #6615
Comet SQL tests (CometSqlFileTestSuite, with fixtures under spark/src/test/resources/sql-tests/) can now express most of what the Scala expression suites check:
per-file Config / ConfigMatrix for ANSI, time zones, dictionary encoding, shuffle mode, batch size and allowIncompatible;
MinSparkVersion / MaxSparkVersion gates;
query modes for native coverage, fallback reasons, errors, tolerance, and native versus codegen dispatch (expect_native / expect_dispatch).
Every query is compared against Spark, and constant folding is disabled so all-literal queries run in Comet.
Expression coverage is still split between those fixtures and the Scala suites. About 1,160 Scala test definitions across 45 suites are in scope here. This epic moves their coverage into SQL tests.
The goal is to lose no coverage, not to translate every test one for one.
A Scala test is deleted once a fixture covers what it covered: the same inputs, configs and assertion strength, or stronger.
Tests that genuinely need Scala stay.
Each sub-issue accounts for every test in its scope, so nothing is dropped silently.
Analysis
I classified every test in the expression suites by static analysis of main at 3bc2faa. That includes the expression-level tests in the aggregate, window and generate operator suites. Each sub-issue lists its tests with a verdict and a target fixture.
The verdicts come from reading the code, not from running fixtures, so treat each target as a starting point. If a conversion turns out not to preserve coverage, keep the Scala test and note why in the sub-issue.
Covered: an existing fixture already covers the test.
Convert: expressible in SQL tests today.
Blocked: needs one of the framework changes below first.
Split: part converts, part stays.
Keep: stays in Scala.
Three findings shape the work.
Few tests are covered today. The fixtures are thinner than the Scala tests:
aggregates are grouped only by string keys;
many fixtures use tolerance= where the Scala test compares exactly;
cast fixtures are non-ANSI only, with no try_cast;
datetime fixtures have far fewer formats, time zones and edge values;
most array and map fixtures run one-row batches.
Most of the work is therefore extending fixtures, not deleting tests.
Some existing tests, Scala and SQL, don't test what they claim. These need fixing as part of the move:
Regex escapes. Spark drops the backslash before an unknown escape, so '\d' is d. Seven regex fixtures, and the Scala rlike tests, never test \d, \w, anchors or backreferences.
NaN under tolerance.query tolerance= skips the comparison when either side is NaN. The NaN results of every per-function math fixture and corr.sql's NaN permutations are never compared.
Literal casts.CometCast folds a cast of a literal on the JVM, so literal casts in the cast fixtures never run native code.
NULL literals. Spark's NullPropagation folds NULL-literal arguments before Comet sees them; the contributor guide's own ascii(NULL) example is affected.
Comet compared with Comet. About 20 cast tests use a helper, assertDataFrameEqualsWithExceptions, that compares Comet with Comet.
Vacuous Scala tests, for example:
three "null group key" aggregate tests build their keys with null.asInstanceOf[Int], which is 0;
a cast test maps over a lazy Iterator and never runs;
try_add / try_subtract / try_multiply are constant-folded by Spark;
the INT96 conversion legs compare Spark with Spark.
About 71 tests need framework changes first.
Structured error parity: expect_error checks only a message substring and never checks that Comet ran the query.
Optimizer-rule control: a fixture can't keep ConstantFolding, which the folded-literal tests need, or exclude NullPropagation / ConvertToLocalRelation.
Ground rules for every PR
Account for every Scala test in scope. Each one is converted, cited as already covered, or kept with its reason. The table in each sub-issue is the checklist.
Delete each converted Scala test in the same PR that adds its coverage. When a suite is empty, delete it and remove it from the suite lists in .github/workflows/pr_build_linux.yml and pr_build_macos.yml. A stale name there doesn't fail CI: CometExpressionCoverageSuite is still listed, although chore: remove coverage file auto generator #2854 deleted it.
Make sure the fixture reaches Comet.
Read expression inputs from Parquet tables, not from inline VALUES, literal casts or NULL literals.
Use one-file inserts (INSERT ... SELECT /*+ COALESCE(1) */ ... or range(0, n, 1, 1)) when batch shape or row order matters.
The docs sub-issue lists all of the traps.
Check special float values exactly. Put NaN, ±Infinity and signed-zero results in a plain query, and write negative zero as double('-0.0').
Keep mechanism checks. Where the Scala test pinned native versus dispatch, pin it with expect_native / expect_dispatch.
Random-data tests convert to curated values that include the generator's edge values. The fuzz suites keep the random coverage: CometFuzzTestSuite, CometFuzzAggregateSuite, CometFuzzMathSuite and CometCodegenFuzzSuite.
Spot-check regression tests named after an issue: confirm the fixture fails with the fix reverted.
Run every profile. Pull request CI runs only the default Spark profile, and fixtures carry version gates, so these PRs need the run-all-spark-profiles label.
Watch CI bucket balance. Moving aggregate and window coverage shifts CI time from the exec bucket to the expressions bucket. Rebalance if expressions grows much longer than the others.
What stays in Scala
Plan shape: operator counts, canonicalization, exchange reuse, mixed-engine partial/final safety.
Metrics, and explain / COMET-INFO text.
Serde and proto internals, and unit tests of helpers such as the CometRegex analyzer and the time-zone ID mapping.
Comet-only failures that intentionally differ from Spark.
Work
Suggested order:
Start with the framework and docs issues.
The tolerance fix is small, and it changes how the math and aggregate conversions are written.
The error-parity and optimizer-rule issues unblock the tests counted as blocked in the table.
The area issues are independent of each other and can proceed in parallel. Window, hash/bitwise and generators are the most mechanical. Float semantics is the lowest priority.
What / Why
Comet SQL tests (
CometSqlFileTestSuite, with fixtures underspark/src/test/resources/sql-tests/) can now express most of what the Scala expression suites check:Config/ConfigMatrixfor ANSI, time zones, dictionary encoding, shuffle mode, batch size and allowIncompatible;MinSparkVersion/MaxSparkVersiongates;expect_native/expect_dispatch).Every query is compared against Spark, and constant folding is disabled so all-literal queries run in Comet.
Expression coverage is still split between those fixtures and the Scala suites. About 1,160 Scala test definitions across 45 suites are in scope here. This epic moves their coverage into SQL tests.
The goal is to lose no coverage, not to translate every test one for one.
Analysis
I classified every test in the expression suites by static analysis of
mainat 3bc2faa. That includes the expression-level tests in the aggregate, window and generate operator suites. Each sub-issue lists its tests with a verdict and a target fixture.The verdicts come from reading the code, not from running fixtures, so treat each target as a starting point. If a conversion turns out not to preserve coverage, keep the Scala test and note why in the sub-issue.
Three findings shape the work.
Few tests are covered today. The fixtures are thinner than the Scala tests:
tolerance=where the Scala test compares exactly;try_cast;Most of the work is therefore extending fixtures, not deleting tests.
Some existing tests, Scala and SQL, don't test what they claim. These need fixing as part of the move:
'\d'isd. Seven regex fixtures, and the Scalarliketests, never test\d,\w, anchors or backreferences.query tolerance=skips the comparison when either side is NaN. The NaN results of every per-function math fixture andcorr.sql's NaN permutations are never compared.CometCastfolds a cast of a literal on the JVM, so literal casts in the cast fixtures never run native code.NullPropagationfolds NULL-literal arguments before Comet sees them; the contributor guide's ownascii(NULL)example is affected.assertDataFrameEqualsWithExceptions, that compares Comet with Comet.null.asInstanceOf[Int], which is0;Iteratorand never runs;try_add/try_subtract/try_multiplyare constant-folded by Spark;About 71 tests need framework changes first.
expect_errorchecks only a message substring and never checks that Comet ran the query.NullPropagation/ConvertToLocalRelation.Ground rules for every PR
.github/workflows/pr_build_linux.ymlandpr_build_macos.yml. A stale name there doesn't fail CI:CometExpressionCoverageSuiteis still listed, although chore: remove coverage file auto generator #2854 deleted it.VALUES, literal casts or NULL literals.INSERT ... SELECT /*+ COALESCE(1) */ ...orrange(0, n, 1, 1)) when batch shape or row order matters.query, and write negative zero asdouble('-0.0').expect_native/expect_dispatch.CometFuzzTestSuite,CometFuzzAggregateSuite,CometFuzzMathSuiteandCometCodegenFuzzSuite.run-all-spark-profileslabel.execbucket to theexpressionsbucket. Rebalance ifexpressionsgrows much longer than the others.What stays in Scala
CometRegexanalyzer and the time-zone ID mapping.InSet.Unevaluableexpressions such asDays/Hours.Work
Suggested order:
Framework and docs (do first)
query tolerance=in Comet SQL tests passes when either side is NaN #6616 test:query tolerance=in Comet SQL tests passes when either side is NaNBy area
Related:
routing_*fixtures are its home.