Describe the bug
Since #5681, the native Parquet scan checks the type conversion for every nested field when it opens a file, while Spark checks it only when it decodes a value. So a file with nothing to decode, either empty or with every row group pruned, whose nested field has a type that the read schema can't convert, now fails the query with SchemaColumnConvertNotSupportedException. Spark returns the rows from the other files, and so did 1.0.0.
When the file does have rows to decode, 1.1.0 now fails the way Spark does. That part is a fix: 1.0.0 returned null for the value it couldn't convert.
Steps to reproduce
withTempPath { dir =>
val path = dir.getCanonicalPath
spark.sql("select named_struct('x', 1) as s").write.parquet(path)
// An empty file whose nested field is a string
spark.sql("select named_struct('x', 'a') as s where false").write.mode("append").parquet(path)
checkSparkAnswer(spark.read.schema("s struct<x:int>").parquet(path))
}
Spark and 1.0.0 return [[1]]. 1.1.0-rc1 fails when it opens the empty file, for column [s, x], BINARY to int. A file whose only row group a filter prunes fails the same way, for example a nested decimal(10,4) read as decimal(10,2) with WHERE id = 100.
Expected behavior
The rows Spark returns.
Additional context
Found by the 1.1.0 regression audit (#6399) and tracked in #6402.
Describe the bug
Since #5681, the native Parquet scan checks the type conversion for every nested field when it opens a file, while Spark checks it only when it decodes a value. So a file with nothing to decode, either empty or with every row group pruned, whose nested field has a type that the read schema can't convert, now fails the query with
SchemaColumnConvertNotSupportedException. Spark returns the rows from the other files, and so did 1.0.0.When the file does have rows to decode, 1.1.0 now fails the way Spark does. That part is a fix: 1.0.0 returned null for the value it couldn't convert.
Steps to reproduce
withTempPath { dir => val path = dir.getCanonicalPath spark.sql("select named_struct('x', 1) as s").write.parquet(path) // An empty file whose nested field is a string spark.sql("select named_struct('x', 'a') as s where false").write.mode("append").parquet(path) checkSparkAnswer(spark.read.schema("s struct<x:int>").parquet(path)) }Spark and 1.0.0 return
[[1]]. 1.1.0-rc1 fails when it opens the empty file, for column[s, x], BINARY to int. A file whose only row group a filter prunes fails the same way, for example a nesteddecimal(10,4)read asdecimal(10,2)withWHERE id = 100.Expected behavior
The rows Spark returns.
Additional context
Found by the 1.1.0 regression audit (#6399) and tracked in #6402.