Skip to content

test: compare in-memory caches with independent answers - #6262

Merged
andygrove merged 3 commits into
apache:mainfrom
LinSimon-901101:fix/6203-cache-test-oracles
Sep 29, 2026
Merged

andygrove merged 3 commits into
apache:mainfrom
LinSimon-901101:fix/6203-cache-test-oracles

Conversation

@LinSimon-901101

Copy link
Copy Markdown
Contributor

Which issue does this PR close?

Closes #6203.

Rationale for this change

Disabling Comet does not bypass a cached relation or its static serializer. Comparing a cached query with checkSparkAnswer can therefore accept the same corrupted values or incorrect pruning on both sides. These tests need expected answers computed independently of the cache.

What changes are included in this PR?

  • Replace cached-query self-comparisons with explicit expected rows or Spark results collected before caching, including Kryo round trips and non-Arrow columnar inputs.
  • Exercise NaN equality/range predicates that use cached statistics and assert the number of retained batches.
  • Add differential tests for native Arrow, Spark columnar, and row input. Compare full values and 23 selective predicates against the original in-memory fixture, preserving independent signed-zero ground truth. Verify writer plans, small batches, and exact batch-pruning counts.
  • Register the new suite in Linux and macOS CI.

How are these changes tested?

  • Rebuilt the native library and ran CometInMemoryCacheSuite, CometInMemoryCacheKryoSuite, and CometInMemoryCachePruningSuite locally across Spark 3.4–4.2: Spark 3.4 had 56 passed / 5 version-gated cancellations; Spark 3.5 had 59 / 2; Spark 4.0, 4.1, and 4.2 had 61 passed each. Spark 4.0 also passed semantic Scalafix CHECK.
  • Fault injection at commit 4924c36d3: omitting NaN from double statistics makes the NaN test and all three differential writer tests fail (4/4). Adding 1.0 to cached doubles causes 9 failures, including non-Arrow columnar input, both Kryo storage levels, and all three new writer tests. Production source was restored afterward and all 61 Spark 4.1 cache tests passed again.
  • Fork CI at the submitted commit passed, including Required Checks, native build, Rust tests, all four Linux Spark 4.1 / JDK 17 test groups, static checks, TPC-H, and TPC-DS under three join configurations. The exec group passed 1,079 tests with 0 failures and 0 cancellations; its logs confirm all 61 cache tests ran, including the three new writer cases. Optional workflows followed the repository's normal PR conditions.

@github-actions github-actions Bot added enhancement New feature or request test Testing related labels Sep 27, 2026

@sunchao sunchao left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Summary

  • Prior state and problem: checkSparkAnswer could compare a cached query against the same cached payload and statistics, allowing incorrect values or pruning to pass.
  • Design approach: Replace those comparisons with explicit rows or Spark answers collected before caching. Add a differential pruning suite for native Arrow, Spark columnar, and row inputs.
  • Correctness / compatibility analysis: Reviewed all three PR commits and all five changed files, including surrounding cache and test-helper code. Compared relevant Spark sources for 3.4.3, 3.5.9, 4.0.4, 4.1.3, and 4.2.0. The NaN comparisons, signed-zero fixture, pruning counts, and independent expected answers match the applicable semantics.
  • Key design decisions: Keeping the pruning oracle in memory avoids Parquet pruning influencing signed-zero expectations. Writer-plan and batch-count assertions verify that the intended paths execute. The local helper and parameterized suite keep the design straightforward.
  • Implementation sketch: Update existing cache and Kryo assertions, add 23 predicates across three writer paths, and register the new suite in both Linux and macOS workflows.
  • Behavioral changes worth calling out: Production behavior is unchanged. Tests now detect cache errors independently. The three new cases took approximately 2.2 seconds combined in the verified fork CI run.
  • Suggested improvements: None at P1/P2 severity. No introduced P1/P2 issues found within this review.

Reviewed full SHA: 4924c36d3e4655909349bb11238905aed1957179. Scope was the full PR diff against base 74fdec5bcdb95f209104a2b698712f668aea4ae2, using merge base f7952de7313775f78ba41eebca2020c4915c16ae. The PR was not a draft. The snapshot and live discussion contained no reviews, issue comments, inline comments, or review threads.

Routed skills: review-comet-pr, review-comet-expression-pr, and review-comet-ffi-pr.

Exact-head CI: Upstream has a successful label check but no build/test verdict. Independently inspected fork run 36267399368: Required Checks, native build, Rust tests, and all four Linux Spark 4.1/JDK 17 groups passed. Its tested merge commit has the same tree as the reviewed head. Logs confirm all 61 cache tests passed within the exec group's 1,079 successful tests.

Validation limits: Local dev/ci/check-ci-config.py and git diff --check passed. JVM/native suites were not rerun locally because this checkout lacks Maven dependencies and a built native library. macOS, nondefault Spark runtime profiles, Spark SQL, and Iceberg suites were skipped in the verified fork run. The author's cross-version runs and fault-injection results were not independently reproduced.

@andygrove andygrove left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I checked this by planting the two bugs from #6203 in ArrowCachedBatchSerializer. With both in, 10 of the 61 cache tests fail, where 3 of 58 failed before. An int corruption confined to the Spark-vectorized convert path used to pass everything and now fails the pruning suite's Spark columnar case. Each writer's corruption is caught by its own case. The three suites pass on 3.4, 3.5, 4.0, 4.1 and 4.2.

"dec >= -1.125 AND dec < 2.125",
"ts < TIMESTAMP '1960-01-05 00:00:00'",
"b <=> true",
"id IN (1, 9, 49)",

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

No predicate here can catch an off-by-one bound on an int column. id IN (1, 9, 49) lands on the second row of each of its batches, and n only appears in n IS NULL, which prunes on the null count. If every int batch reports its upper bound one lower, this suite still passes on all three writers, and only the typed-bounds tests in CometInMemoryCacheSuite notice. Could we add n = 5? n is constant within a batch, so that one predicate sits on both bounds of batch 5.

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

@andygrove Thanks for pointing this out. I'll add the n = 5 coverage in a follow-up issue.

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Opened #6378 to track adding n = 5 coverage for integer bounds.

@andygrove
andygrove added this pull request to the merge queue Sep 28, 2026
Merged via the queue into apache:main with commit ba9aa33 Sep 29, 2026
38 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

enhancement New feature or request test Testing related

Projects

None yet

Development

Successfully merging this pull request may close these issues.

In-memory cache tests that use checkSparkAnswer compare the cache with itself

3 participants