Describe the bug
When a field inside a struct column, or inside the struct of an array, is renamed with ALTER TABLE ... RENAME COLUMN, the native Iceberg scan returns NULL for that field in rows from data files written before the rename. There is no error, and Spark returns the values.
This is a regression in 1.1.0. 1.0.0 returns the right values for the queries below.
iceberg-rust compares the file's nested type with the table's using arrow's equals_datatype, which ignores field names. When the two match, it passes the column through. The batch is labeled with the table's field names, but the struct array keeps the file's old names (apache/iceberg-rust#2617). In 1.0.0, CometCastColumnExpr relabeled such a column by position to the names Spark asked for. In 1.1.0 the column goes through DataFusion's struct cast instead, which matches fields by the array's names and fills the renamed field with NULL. This most likely dates from #5262, which added is_pure_structural_narrowing to keep DataFusion's CastExpr when the declared struct field names match.
The bug shows up when the file's nested fields have the same nullability as the table's, which is the usual case. Iceberg writes a file with the nullability of the inserted data, so any null inside the struct makes the fields optional, as the table declares them. Spark 4 writes a struct built only from non-null literals with required fields. iceberg-rust then casts the column by position instead, and a plain rename reads correctly. That is why the reproduction below includes a row with nulls. Spark 3.4 writes the fields as optional either way.
Steps to reproduce
CREATE TABLE cat.db.evo_rename (id INT, s STRUCT<a: INT, b: STRING>, items ARRAY<STRUCT<a: INT, b: INT>>) USING iceberg;
INSERT INTO cat.db.evo_rename VALUES
(1, named_struct('a', 1, 'b', 'x'), array(named_struct('a', 1, 'b', 2))),
(2, named_struct('a', CAST(NULL AS INT), 'b', CAST(NULL AS STRING)), array(CAST(NULL AS STRUCT<a: INT, b: INT>)));
ALTER TABLE cat.db.evo_rename RENAME COLUMN s.a TO z;
ALTER TABLE cat.db.evo_rename RENAME COLUMN items.element.a TO z;
SELECT id, s FROM cat.db.evo_rename ORDER BY id;
SELECT id, items FROM cat.db.evo_rename ORDER BY id;
With the native Iceberg scan on, which is the default, 1.1.0-rc1 returns (1, {null, x}) for the first query and (1, [{null, 2}]) for the second.
Expected behavior
Spark's results, which 1.0.0 also returns: (1, {1, x}) and (2, {null, null}) for the first query, and (1, [{1, 2}]) and (2, [null]) for the second.
Workaround
spark.comet.scan.icebergNative.enabled=false reads every Iceberg table with Spark's reader.
Additional context
Two related cases were already wrong in 1.0.0:
- Reading only the renamed field,
SELECT id, s.z, returned NULL in 1.0.0. In 1.1.0 it fails with Cannot cast struct with 2 fields to 1 fields because there is no field name overlap.
- A reorder plus a rename,
ALTER COLUMN m.b FIRST followed by RENAME COLUMN m.a TO c on m STRUCT<a: INT, b: INT>, returns {1, 2} in 1.0.0 and {2, null} in 1.1.0, where Spark returns {2, 1}.
Verified on Spark 4.1 with Iceberg 1.11.0, comparing the 1.0.0 tag (3a7a2c4) with 1.1.0-rc1 (470fc78). The same NULLs appear on main on Spark 3.4 and 4.1.
Found while fixing #6504. #6543 makes the native scan fall back when a projected column has a nested field that was added or renamed over the table's schema history, so it fixes this too. apache/iceberg-rust#3255 reconciles nested fields by field id, which would fix the iceberg-rust side.
Describe the bug
When a field inside a struct column, or inside the struct of an array, is renamed with
ALTER TABLE ... RENAME COLUMN, the native Iceberg scan returns NULL for that field in rows from data files written before the rename. There is no error, and Spark returns the values.This is a regression in 1.1.0. 1.0.0 returns the right values for the queries below.
iceberg-rust compares the file's nested type with the table's using arrow's
equals_datatype, which ignores field names. When the two match, it passes the column through. The batch is labeled with the table's field names, but the struct array keeps the file's old names (apache/iceberg-rust#2617). In 1.0.0,CometCastColumnExprrelabeled such a column by position to the names Spark asked for. In 1.1.0 the column goes through DataFusion's struct cast instead, which matches fields by the array's names and fills the renamed field with NULL. This most likely dates from #5262, which addedis_pure_structural_narrowingto keep DataFusion'sCastExprwhen the declared struct field names match.The bug shows up when the file's nested fields have the same nullability as the table's, which is the usual case. Iceberg writes a file with the nullability of the inserted data, so any null inside the struct makes the fields optional, as the table declares them. Spark 4 writes a struct built only from non-null literals with required fields. iceberg-rust then casts the column by position instead, and a plain rename reads correctly. That is why the reproduction below includes a row with nulls. Spark 3.4 writes the fields as optional either way.
Steps to reproduce
With the native Iceberg scan on, which is the default, 1.1.0-rc1 returns
(1, {null, x})for the first query and(1, [{null, 2}])for the second.Expected behavior
Spark's results, which 1.0.0 also returns:
(1, {1, x})and(2, {null, null})for the first query, and(1, [{1, 2}])and(2, [null])for the second.Workaround
spark.comet.scan.icebergNative.enabled=falsereads every Iceberg table with Spark's reader.Additional context
Two related cases were already wrong in 1.0.0:
SELECT id, s.z, returned NULL in 1.0.0. In 1.1.0 it fails withCannot cast struct with 2 fields to 1 fields because there is no field name overlap.ALTER COLUMN m.b FIRSTfollowed byRENAME COLUMN m.a TO conm STRUCT<a: INT, b: INT>, returns{1, 2}in 1.0.0 and{2, null}in 1.1.0, where Spark returns{2, 1}.Verified on Spark 4.1 with Iceberg 1.11.0, comparing the 1.0.0 tag (3a7a2c4) with 1.1.0-rc1 (470fc78). The same NULLs appear on
mainon Spark 3.4 and 4.1.Found while fixing #6504. #6543 makes the native scan fall back when a projected column has a nested field that was added or renamed over the table's schema history, so it fixes this too. apache/iceberg-rust#3255 reconciles nested fields by field id, which would fix the iceberg-rust side.