Skip to content

[SPARK-49110][SQL] Simplify SubqueryAlias.metadataOutput to always propagate metadata columns - #53861

Closed
cloud-fan wants to merge 2 commits into
apache:masterfrom
cloud-fan:meta_col
Closed

cloud-fan wants to merge 2 commits into
apache:masterfrom
cloud-fan:meta_col

Conversation

@cloud-fan

Copy link
Copy Markdown
Contributor

What changes were proposed in this pull request?

This PR simplifies SubqueryAlias.metadataOutput to always propagate metadata columns from its child, rather than only propagating when the child is a LeafNode or another SubqueryAlias.

The previous implementation was introduced in SPARK-40149 as a workaround to forbid queries like SELECT m FROM (SELECT a FROM t) while still allowing DataFrame API chaining. However, this created an inconsistency since SubqueryAlias should conceptually just rename/qualify columns, not filter which ones are accessible.

With this change:

  • SubqueryAlias always propagates metadataOutput (with qualifier applied)
  • The qualifiedAccessOnly filter is preserved to handle natural join metadata columns
  • Queries like SELECT m FROM (SELECT a FROM t) AS alias now work, consistent with how Project already propagates metadata columns

Why are the changes needed?

  1. Consistency: SubqueryAlias is a rename operation and should not selectively block metadata column propagation
  2. Simpler code: Removes the special-case logic checking for LeafNode/SubqueryAlias children
  3. Better error messages: When metadata columns from both sides of a join have the same name, users now get an "ambiguous reference" error rather than "column not found"

Does this PR introduce any user-facing change?

Yes, queries that previously failed with "column not found" when accessing metadata columns through a subquery alias will now succeed (if unambiguous) or fail with "ambiguous reference" (if multiple columns have the same name).

How was this patch tested?

Updated existing tests and added new test for ambiguous metadata columns after join with SubqueryAlias.

Was this patch authored or co-authored using generative AI tooling?

Yes.

@github-actions

github-actions Bot commented Jan 20, 2026 •

Copy link
Copy Markdown

JIRA Issue Information

=== Improvement SPARK-49110 ===
Summary: Unable to access _metadata column of tables with CHAR column with reader side padding enabled
Assignee: None
Status: Open
Affected: ["3.5.1","4.0.0"]


This comment was automatically generated by GitHub Actions

@github-actions github-actions Bot added the SQL label Jan 20, 2026
@cloud-fan
cloud-fan marked this pull request as draft January 20, 2026 01:56
@cloud-fan cloud-fan changed the title [SPARK-XXXXX][SQL] Simplify SubqueryAlias.metadataOutput to always propagate metadata columns [SPARK-49110][SQL] Simplify SubqueryAlias.metadataOutput to always propagate metadata columns Jan 20, 2026
@cloud-fan
cloud-fan marked this pull request as ready for review January 20, 2026 05:59
@cloud-fan
cloud-fan force-pushed the meta_col branch 2 times, most recently from 3ff52dd to 4921f9b Compare January 20, 2026 06:10
…opagate metadata columns

### What changes were proposed in this pull request?

This PR simplifies `SubqueryAlias.metadataOutput` to always propagate metadata columns from its child, rather than only propagating when the child is a `LeafNode` or another `SubqueryAlias`.

The previous implementation was introduced in SPARK-40149 as a workaround to forbid queries like `SELECT m FROM (SELECT a FROM t)` while still allowing DataFrame API chaining. However, this created an inconsistency since `SubqueryAlias` should conceptually just rename/qualify columns, not filter which ones are accessible.

With this change:
- `SubqueryAlias` always propagates `metadataOutput` (with qualifier applied)
- The `qualifiedAccessOnly` filter is preserved to handle natural join metadata columns
- Queries like `SELECT m FROM (SELECT a FROM t) AS alias` now work, consistent with how `Project` already propagates metadata columns

### Why are the changes needed?

1. **Consistency**: `SubqueryAlias` is a rename operation and should not selectively block metadata column propagation
2. **Simpler code**: Removes the special-case logic checking for `LeafNode`/`SubqueryAlias` children
3. **Better error messages**: When metadata columns from both sides of a join have the same name, users now get an "ambiguous reference" error rather than "column not found"

### Does this PR introduce _any_ user-facing change?

Yes, queries that previously failed with "column not found" when accessing metadata columns through a subquery alias will now succeed (if unambiguous) or fail with "ambiguous reference" (if multiple columns have the same name).

### How was this patch tested?

Updated existing tests and added new test for ambiguous metadata columns after join with SubqueryAlias.
Comment thread sql/catalyst/src/main/scala/org/apache/spark/sql/internal/SQLConf.scala Outdated
@cloud-fan

Copy link
Copy Markdown
Contributor Author

thanks for the review, merging to master!

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants