Describe the bug
Spark treats two struct types whose field names differ only in case as the same type when spark.sql.caseSensitive is false, the default. So it accepts CASE WHEN branches like named_struct('x', i) and named_struct('X', i) without a cast. CaseWhen.dataType merges the branch types left to right, starting from the first THEN branch, so the result's fields carry the first THEN branch's names.
Since #6350, native CASE WHEN takes its common type from DataFusion's get_coerce_type_for_case_expression, which folds starting from the ELSE branch. It then casts every branch to that type, so the native result carries the ELSE branch's field names on every row.
The values are right, and checkSparkAnswer passes, because row comparison ignores struct field names. The names show up wherever native code reads them from the Arrow type. One example is native to_json, which is opt-in (spark.comet.expression.StructsToJson.allowIncompatible=true). With the default configs, to_json runs through the codegen dispatcher, which evaluates the CASE WHEN in the JVM, so the output is right.
This is on main only, because #6350 isn't in 1.1.0.
Steps to reproduce
// In a suite extending CometTestBase, on main (9c7fcc5aa4)
withSQLConf("spark.comet.expression.StructsToJson.allowIncompatible" -> "true") {
sql("CREATE TABLE t(q boolean, i int) USING parquet")
sql("INSERT INTO t SELECT id % 2 = 0, CAST(id AS INT) FROM range(0, 4, 1, 1)")
checkSparkAnswer(
"SELECT to_json(CASE WHEN q THEN named_struct('x', i) ELSE named_struct('X', i) END) FROM t")
}
Expected behavior
Spark returns {"x":0}, {"x":1}, {"x":2} and {"x":3}. Comet, with the CASE WHEN and to_json in a CometProject, returns {"X":0}, {"X":1}, {"X":2} and {"X":3}.
Additional context
#6458 builds native IF through create_if_expr, which runs the same coercion starting from the THEN branch for this reason. Folding from the first THEN branch in create_case_when too, with the ELSE branch last, would match CaseWhen.dataType.
Found while addressing review on #6458.
Describe the bug
Spark treats two struct types whose field names differ only in case as the same type when
spark.sql.caseSensitiveis false, the default. So it acceptsCASE WHENbranches likenamed_struct('x', i)andnamed_struct('X', i)without a cast.CaseWhen.dataTypemerges the branch types left to right, starting from the first THEN branch, so the result's fields carry the first THEN branch's names.Since #6350, native
CASE WHENtakes its common type from DataFusion'sget_coerce_type_for_case_expression, which folds starting from the ELSE branch. It then casts every branch to that type, so the native result carries the ELSE branch's field names on every row.The values are right, and
checkSparkAnswerpasses, because row comparison ignores struct field names. The names show up wherever native code reads them from the Arrow type. One example is nativeto_json, which is opt-in (spark.comet.expression.StructsToJson.allowIncompatible=true). With the default configs,to_jsonruns through the codegen dispatcher, which evaluates theCASE WHENin the JVM, so the output is right.This is on
mainonly, because #6350 isn't in 1.1.0.Steps to reproduce
Expected behavior
Spark returns
{"x":0},{"x":1},{"x":2}and{"x":3}. Comet, with theCASE WHENandto_jsonin aCometProject, returns{"X":0},{"X":1},{"X":2}and{"X":3}.Additional context
#6458 builds native
IFthroughcreate_if_expr, which runs the same coercion starting from the THEN branch for this reason. Folding from the first THEN branch increate_case_whentoo, with the ELSE branch last, would matchCaseWhen.dataType.Found while addressing review on #6458.