Skip to content

[VL] Struct fields are read from the wrong column when two same-named struct columns meet in one projection #13197

Description

@LuciferYang

Backend

Velox (Bolt shares the same binding code).

Bug description

When a projection's input holds two struct attributes with the same name but different field layouts, Gluten reads struct fields from the wrong one. ExpressionConverter.bindGetStructField re-resolves the root attribute of a GetStructField chain by name over the whole input and keeps the last match, so l.s.x and r.s.x are both bound against whichever s comes last.

Repro with the default configuration:

spark.sql("select id, named_struct('x', id, 'y', id * 100) as s from range(5)").write.parquet("/tmp/l")
spark.sql("select id, named_struct('y', id * 100, 'x', id) as s from range(5)").write.parquet("/tmp/r")
spark.read.parquet("/tmp/l").createOrReplaceTempView("l")
spark.read.parquet("/tmp/r").createOrReplaceTempView("r")
spark.sql("select l.s.x, l.s.y, r.s.x, r.s.y from l join r on l.id = r.id").show()

Vanilla Spark returns [1, 100, 1, 100] for id 1. Gluten returns [100, 1, 1, 100].

The same function also looks each field up by name and keeps the last match, which misreads structs with duplicate field names. The Velox generate rewrite builds such a struct from the generator output aliases, so select v.* from (select id, array(named_struct('p', id, 'q', id * 10)) as arr from range(5)) lateral view inline(arr) v as x, x returns [10, 10] for id 1 instead of [1, 10].

Gluten version

main (ca7a0a4)

Spark version

Spark 3.5

This issue was written with the assistance of AI (Claude Opus).

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions