Repository navigation
Conversation
1bf4f08 to
6ad5c0d
Compare
|
Status updated 2026-10-07: #6350, #6396 and #6397 are merged. Feature refresh Historical integration progress and validation#6350 has merged. #6396 remains open with passing PR CI and is the remaining implementation prerequisite. #6397 contributes independent IF semantic tests and does not block this feature. The refreshed integration candidate Local validation of this exact tree passed: release native build; 24 Spark 4.1.3 tests; six Spark 3.5 tests with strict warnings (both on JDK 21); 16 conditional, seven AtLeastNNonNulls and 50 planner Rust tests; workspace/all-target release Clippy, cargo fmt, Spotless and RAT. Neither Spark run had failures, cancellations or skipped tests. Two independent processes passed all full-query answers and native/dispatcher-path assertions on the M3 Pro/Spark 4.1.3 workload: 1,048,576 Parquet rows, 32 Double columns, no NULLs, The refreshed candidate is published in fork draft #42. Current-head fork CI passed at Upstream #6180 remains Draft at |
| } | ||
| } | ||
|
|
||
| test("filter output projection preserves rows, aliases and metrics") { |
There was a problem hiding this comment.
Let's see if you could port this test to Comet SQL tests: https://datafusion.apache.org/comet/contributor-guide/sql-file-tests.html
There was a problem hiding this comment.
Moved the projection result regression to spark/src/test/resources/sql-tests/operators/filter_projection.sql in #6396, which is now merged. It covers count-only output, aliases/order/duplicates, full projections, empty results and a rejected row that must not reach an ANSI cast. The Scala test retains the native output-row metrics checks, which the SQL fixture cannot assert. These generic projection changes are no longer in this PR’s diff.
| }) | ||
| .collect(); | ||
|
|
||
| let mut exprs = ProjectionExprs::from(exprs?); |
There was a problem hiding this comment.
Is it possible to fix this issue in a separate PR?
There was a problem hiding this comment.
| result, None, | ||
| )))); | ||
| } | ||
| if batch.num_rows() < 64 { |
There was a problem hiding this comment.
64 was a one-word starting point for the general counting path: bitmap counters process 64 rows per word, while small batches may not amortize their bookkeeping. It is a heuristic, not a required or universal crossover. I made it configurable as spark.comet.exec.atLeastNNonNulls.smallBatchThreshold, default 64; any positive value is accepted, independently of the bitmap word size and the na.drop threshold. Tests cover both strategies and non-word-aligned thresholds. Criterion now compares row, bitmap and default strategies on identical inputs around the boundary. Current-head timings are still pending, so I am not claiming that 64 is optimal.
Which issue does this PR close?
Closes #6093.
Rationale for this change
Spark uses
AtLeastNNonNullsforDataFrame.na.drop. Native evaluation avoids falling back to Spark while preserving the count of non-NULL, non-NaN values and Spark's per-row short-circuit and error behavior.What changes are included in this PR?
spark.comet.exec.atLeastNNonNulls.smallBatchThreshold, default 64. It selects the general counting strategy by input batch row count and accepts any positive integer. It changes neither the fixed 64-bit word size nor thena.dropthreshold.Shared CASE/IF evaluation (#6350), filter-output pruning (#6396), and independent IF semantic tests (#6397) have merged. The branch uses these from main. Generic projection code is no longer part of this PR's diff. The projection result regression is in
operators/filter_projection.sql; its Scala test retains native output-row metrics checks.How are these changes tested?
Current published head:
1f0c6478ecde67d31463bb23249f334bc26fa35c. Current-head upstream CI completed successfully: 25 checks passed, with 15 conditional jobs skipped.at_least_n_non_nullsunit tests and the planner threshold round-trip test. The workspace run completed with 2,228 passing tests and 5 skipped tests.AtLeastNNonNullsJVM semantic tests and the query-serde threshold test. The job completed with 2,232 passing tests, 1 cancellation and 12 ignored tests, with no failures or aborted suites.Historical combined-candidate tests and timings remain in fork #42 at
4528a2a67cff87c06179337ce4e1b35de2355839and its completed CI. They do not establish current-head performance. The original issue author's workload remains unverified.Keep this PR Draft pending the remaining Ready requirements: relevant Spark SQL/other-profile CI and current-head benchmark controls. Separate count-only output pruning from full-column conditional-expression costs, compare both counting strategies on identical inputs, and report repeated-run variation and slower cases. Default CI establishes the native/JVM test verdict; it does not provide those broader suite or performance results.