Skip to content

Native IF fails with "column types must match schema types" when its branches differ in nested nullability #6334

Description

@andygrove

Describe the bug

Comet's native IF fails at runtime when its two branches have the same Spark type but differ in nested nullability, for example a struct field or map value that can be NULL in one branch and not in the other:

org.apache.comet.CometNativeException: Invalid argument error: column types must match schema types, expected Struct("x": Int32) but found Struct("x": non-null Int32) at column index 0

Spark treats the two branch types as the same when only their nullability differs, so it adds no cast, and Comet serializes each branch with its own Arrow type. The native IfExpr reports the THEN branch's type as its output type. When every row of a batch takes the ELSE branch, though, it returns the ELSE branch's array unchanged, and that array has the other nullability.

CASE WHEN does not fail, because the planner casts each branch to a common type first.

Steps to reproduce

CREATE TABLE t(q boolean, i int) USING parquet;
INSERT INTO t VALUES (true, 1), (false, 2);

-- both fail
SELECT IF(q, named_struct('x', i), named_struct('x', 0)) FROM t;
SELECT IF(q, map('k', i), map('k', 0)) FROM t;

-- passes
SELECT CASE WHEN q THEN named_struct('x', i) ELSE named_struct('x', 0) END FROM t;

Reproduced on main at 7649361 through CometSqlFileTestSuite. The failure needs a batch in which no row takes the THEN branch, which this two-row table gives in one of its tasks. IF(q, array(i), array(0)) passes.

Expected behavior

The same result as Spark.

Additional context

Found while adding tests for #3025. The planner could build IF the way it builds CASE WHEN, casting a branch whose Arrow type differs from the common type.

Activity

  1. added
    regressionA bug that did not affect the most recent Comet release
    on Sep 29, 2026
  2. andygrove commented on Sep 29, 2026

    @andygrove
    MemberAuthor

    This is a regression in 1.1.0. In 1.0.0 both queries above ran in Spark. Comet declined the map and struct literals and fell back for the whole Project. #5452 made complex literals native, and that is what exposes IF to this.

    The 1.1.0 regression audit (#6399) confirmed it on 1.1.0-rc1 with a map column, using IF(c, m, map('z', 0)) and IF(m IS NULL, map('z', 0), m). Both fail on any batch where the predicate is the same for every row, while 1.0.0 returns Spark's answer. I've added the regression label, and it's tracked in #6402.

  3. self-assigned this
    on Oct 1, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

Labels

bugSomething isn't workingregressionA bug that did not affect the most recent Comet releaserequires-triage

Type

No type

Projects

No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions