Describe the bug
Native CASE WHEN and COALESCE compute their common branch type through get_coerce_type_for_case_expression, which folds DataFusion's type_union_coercion. DataFusion reconciles struct fields by name. Spark reconciles them by position and keeps the left branch's field names.
When case-insensitive Spark resolution accepts case-distinct names in different positions, DataFusion can pair fields from different positions and change their data types. For example:
CASE WHEN q
THEN named_struct('x', i, 'X', CAST(5.5 AS DOUBLE))
ELSE named_struct('X', 0, 'x', d)
END
For i INT and d DOUBLE, Spark returns STRUCT<x:INT,X:DOUBLE>. The native common-type calculation can instead pair the first INT field with the second DOUBLE field by name and widen x to DOUBLE. Native to_json then exposes the wrong result as {"x":7.0,"X":5.5} instead of {"x":7,"X":5.5}.
COALESCE uses the same create_case_when path and has the same positional mismatch.
Expected behavior
Match Spark's findTypeForComplex: recursively zip struct fields by position, retain the left branch's field names, and combine nullability. Arrays and maps containing structs must follow the same recursive rule.
Suggested fix and tests
PR #6458 added if_common_type for native IF, with positional reconciliation through structs, arrays, and maps. Extend or generalize that helper for the create_case_when path used by both CASE WHEN and COALESCE.
Add regression coverage for:
CASE WHEN and COALESCE with case-distinct field names swapped by position
- schema and value preservation, including
INT not widening to DOUBLE
- both branch orders
- nested structs in arrays and maps
Native to_json with spark.comet.expression.StructsToJson.allowIncompatible=true can observe the field names and numeric representation.
Additional context
Found during review of #6458. Issue #6482 tracks a related, simpler symptom where native CASE WHEN takes field names from the ELSE branch. This issue tracks the broader positional type-coercion problem and also covers COALESCE.
Describe the bug
Native
CASE WHENandCOALESCEcompute their common branch type throughget_coerce_type_for_case_expression, which folds DataFusion'stype_union_coercion. DataFusion reconciles struct fields by name. Spark reconciles them by position and keeps the left branch's field names.When case-insensitive Spark resolution accepts case-distinct names in different positions, DataFusion can pair fields from different positions and change their data types. For example:
For
i INTandd DOUBLE, Spark returnsSTRUCT<x:INT,X:DOUBLE>. The native common-type calculation can instead pair the firstINTfield with the secondDOUBLEfield by name and widenxtoDOUBLE. Nativeto_jsonthen exposes the wrong result as{"x":7.0,"X":5.5}instead of{"x":7,"X":5.5}.COALESCEuses the samecreate_case_whenpath and has the same positional mismatch.Expected behavior
Match Spark's
findTypeForComplex: recursively zip struct fields by position, retain the left branch's field names, and combine nullability. Arrays and maps containing structs must follow the same recursive rule.Suggested fix and tests
PR #6458 added
if_common_typefor nativeIF, with positional reconciliation through structs, arrays, and maps. Extend or generalize that helper for thecreate_case_whenpath used by bothCASE WHENandCOALESCE.Add regression coverage for:
CASE WHENandCOALESCEwith case-distinct field names swapped by positionINTnot widening toDOUBLENative
to_jsonwithspark.comet.expression.StructsToJson.allowIncompatible=truecan observe the field names and numeric representation.Additional context
Found during review of #6458. Issue #6482 tracks a related, simpler symptom where native
CASE WHENtakes field names from the ELSE branch. This issue tracks the broader positional type-coercion problem and also coversCOALESCE.