Skip to content

Native Parquet scan fails on an empty file whose nested field type differs from the read schema #6506

Description

@andygrove

Describe the bug

Since #5681, the native Parquet scan checks the type conversion for every nested field when it opens a file, while Spark checks it only when it decodes a value. So a file with nothing to decode, either empty or with every row group pruned, whose nested field has a type that the read schema can't convert, now fails the query with SchemaColumnConvertNotSupportedException. Spark returns the rows from the other files, and so did 1.0.0.

When the file does have rows to decode, 1.1.0 now fails the way Spark does. That part is a fix: 1.0.0 returned null for the value it couldn't convert.

Steps to reproduce

withTempPath { dir =>
  val path = dir.getCanonicalPath
  spark.sql("select named_struct('x', 1) as s").write.parquet(path)
  // An empty file whose nested field is a string
  spark.sql("select named_struct('x', 'a') as s where false").write.mode("append").parquet(path)
  checkSparkAnswer(spark.read.schema("s struct<x:int>").parquet(path))
}

Spark and 1.0.0 return [[1]]. 1.1.0-rc1 fails when it opens the empty file, for column [s, x], BINARY to int. A file whose only row group a filter prunes fails the same way, for example a nested decimal(10,4) read as decimal(10,2) with WHERE id = 100.

Expected behavior

The rows Spark returns.

Additional context

Found by the 1.1.0 regression audit (#6399) and tracked in #6402.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

Labels

bugSomething isn't workingregressionA bug that did not affect the most recent Comet releaserequires-triage

Type

No type

Projects

No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions