Skip to content

[core-spec] Clarify name resolution for fields, keys, relationships, and metrics - #486

Draft
justint-db wants to merge 2 commits into
apache:mainfrom
justint-db:clarify-name-resolution
Draft

justint-db wants to merge 2 commits into
apache:mainfrom
justint-db:clarify-name-resolution

Conversation

@justint-db

@justint-db justint-db commented Sep 29, 2026 •

Copy link
Copy Markdown

Summary

The spec doesn't say whether a name like amount in the metric expression SUM(orders.amount) refers to a field declared in the orders dataset or to a column of the warehouse table named by its source. Converters currently disagree (see the dev@ thread "what is amount?" started by Damian Waldron). The same ambiguity applies to primary_key, unique_keys, and relationship from_columns / to_columns.

This PR follows the conclusion of that thread, which mirrors SQL views. A dataset's fields are defined in terms of its source columns, and model-level constructs are defined in terms of dataset fields. They never reach through a dataset to its source columns.

Related issues: #462 (metric expressions), #466 (relationship keys).

Changes

  • core-spec/spec.md: new Name Resolution section that sets out:
    • Three levels of names: source columns, dataset fields, and model-scoped metrics.
    • A table of what a name refers to in each place:
      • A field expression refers to a source column of its own dataset.
      • primary_key, unique_keys, from_columns, and to_columns refer to dataset fields.
      • A metric expression refers to a dataset field, which must be written as dataset.field.
    • A worked example where source column names differ from field names (o_totalprice is exposed as total_amount), with examples of invalid references.
    • A tie-break rule for when a field and a source column share a name.
    • Converters that generate SQL against the source tables must replace each field reference with the field's expression.
    • Features that aren't supported yet but may be added later: field-to-field references (lateral column aliases), metric-to-metric references, and datasets built on other datasets.
  • core-spec/spec.md, other sections:
    • The key and relationship descriptions now say "field names" instead of "columns".
    • The Fields and Metrics sections link to the new section.
    • The Datasets example now declares the fields its keys reference.
    • Version History has a new entry.
  • core-spec/expression_language.md: the Name Spaces section now gives the rules for field and metric expressions instead of pointing to a separate document.
  • core-spec/ossie-schema.json, core-spec/spec.yaml: only descriptions and comments changed. The schema itself is unchanged.
  • examples/tpcds_semantic_model.yaml: added the ss_ticket_number field, which the store_sales primary key and unique key already referenced.

Breaking change

Earlier versions didn't say how these names resolve. This PR does, and marks the change Breaking in Version History:

  • Converters that read metric expressions or relationship/key columns as warehouse columns must now resolve them as dataset fields. Per the survey in the dev@ thread, that's Databricks, dbt, NVIDIA, and in part Omni. Converters that target SQL-direct backends should replace each field reference with the field's expression.
  • Metric expressions must now use qualified dataset.field references. Unqualified field names are invalid.
  • A dataset that declares no fields can't declare keys, take part in relationships, or be referenced by metrics.

Converter test fixtures may need to be updated to follow these rules. A follow-up could add compliance fixtures where physical and logical names differ, as JB suggested on the thread.

AI disclosure

Per the ASF Generative Tooling Guidance, this contribution was prepared with AI assistance. All specification decisions and design choices are mine. I have reviewed and verified every change.

…and metrics

Metric expressions, primary_key, unique_keys, and relationship
from_columns/to_columns refer to dataset fields, not columns of the
dataset's source. Field expressions refer to source columns. Metric
expressions must use qualified dataset.field references.

Adds a Name Resolution section to spec.md, updates expression_language.md,
schema descriptions, and spec.yaml comments, and fixes examples whose keys
referenced undeclared fields.

Co-authored-by: Isaac <no-reply@databricks.com>
Use the spec.md wording for primary_key, unique_keys, from_columns, and
to_columns in ossie-schema.json and spec.yaml, and drop the repeated
view/compilation note from expression_language.md.

Co-authored-by: Isaac <no-reply@databricks.com>
@djwaldo

djwaldo commented Sep 30, 2026

Copy link
Copy Markdown
Contributor

The following is the scenario I have been struggling with when creating the ThoughtSpot convertor. This prompted this initial post regarding the nameing conversion.

In ThoughtSpot our most common architecture is two tiered. Where the following is an example.
Level 0 -> Database Table. I.e. table in SF or DBX
Level 1 -> ThoughtSpot Logical Table. This is a one-to-one representation of the underlying DB table. This can capture and define metadata which is leveraged in Layer 2. Importantly the column DISPLAY NAME can be changed to a business friendly name. This is the reference that layer 2 uses.
Level 2 -> ThoughtSpot Semantic Model. This is built from Level 1. This is where metrics are defined. This level is only aware of the DISPLAY NAMEs as defined in Level 1. There is not knowledge of what the actual column in the DB is defined as.

Convertor Logic.
The initial ThoughtSpot convertor works fine for round-trips of ThoughtSpot models. The metric definitions are defined in ThoughtSpot syntax and then model assumes that Level 1 exists in the target state. However this does not create a functional conversion for the path of ThoughtSpot -> OSSIE -> DBX -> OSSIE -> ThoughtSpot. The initial issues is that the metrics were defined with a custom extension of ThoughtSpot. DBX has no way to parse this into DBX metrics.

The solution I was going to land on (which raised this issue) was to

Therefore the Semantic Layers metric definitions are at columns from layer 2. These do not necessarily represent what is defined in the database. I can convert a TS model to OSSIE and then convert the OSSIE model to DBX. However the database columns are never defined anywhere.

To make this lossless should I export two definitions
TS Dialect -> points at the TS layer 2 as that is how our product works
SQL Dialect -> points at the actual DB columns.

@kayemkim kayemkim left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Ran this merged onto current main (edbcf79) with the workflows the changed paths trigger. Everything passes except Omni: test_tpcds_export_matches_expected fails on all five Pythons. The test reads examples/tpcds_semantic_model.yaml directly, and the new ss_ticket_number field now comes out as a visible dimension with its description, where converters/omni/tests/fixtures/tpcds_omni/views/store_sales.view.yaml still has the hidden: true dimension Omni synthesised for the undeclared key. That one entry is the only difference. The PR title check also rejects the [core-spec] prefix, it wants docs(core-spec): ... or Core-spec: ....

One data point for the breaking-change list. #485, merged yesterday for #466, moved the ThoughtSpot to-ossie direction the other way: on current main, a model whose display names differ from the warehouse names comes out with fields customer_id and id but from_columns: [cust_id_fk], to_columns: [cust_pk] and primary_key: [cust_pk], so its relationships and keys no longer name fields of the dataset. Under this text that converter joins the Databricks, dbt and NVIDIA group, and #466 probably wants reopening or a follow-up. For sequencing, #343 lets a dataset-scoped metric reference source columns unqualified, so the table here would want a row for datasets[].metrics when that lands.

The text follows the thread's conclusion closely, and the tie-break sentence plus the invalid-reference examples are the two things a converter author would otherwise have to ask about. LGTM on the spec side.

@isocline

isocline commented Oct 2, 2026

Copy link
Copy Markdown

+1 on the resolution, and thank you for writing it down — "they never reach
through a dataset to its source columns" is the sentence I was hoping to be
able to point an implementation at.

@djwaldo — on the two-dialect question, I think this PR already answers it,
and the answer is that you do not need the second one.

Your three levels line up with the rules here one-for-one:

ThoughtSpot Ossie, under this PR
Level 0 — DB table/column the dataset's source, and what a field expression names
Level 1 — logical table, display names the dataset's fields: name = display name, expression = the DB column
Level 2 — metrics over display names metric expression, written dataset.field

That is the same shape as the worked example in this PR, where o_totalprice
is exposed as total_amount. Level 1 is not a second namespace that Ossie is
missing — it is the dataset field layer. And the rule that "converters
generating SQL against the source tables must replace each field reference
with the field's expression" is exactly the TS → Ossie → DBX hop: the DBX
converter substitutes total_amount back to o_totalprice on its way out,
so the DB column is defined in exactly one place rather than nowhere.

I would be wary of carrying it in two dialects. dialect says which SQL
flavour an expression is written in; it does not say which naming level the
identifiers inside it come from. If a TS dialect entry resolved against Level
2 names and an ANSI_SQL entry against Level 0 names, a consumer choosing a
dialect would silently change name resolution — which is the class of bug
this PR exists to remove, reintroduced one level down. Two dialects of the
same expression should differ only in syntax, never in what the names mean.

One note on the deferred list, specifically field-to-field references, since
our DSL has that feature shipped and the experience may be useful when it
comes back around.

Ours are a dotted logical path resolved against declared fields
(@ref(board.name)). What made them tractable was not the resolution rule
itself but the restriction that came with it: a reference resolves against
declared names only, and an unresolvable path is an authoring-time failure
rather than something that degrades at query time. Once that held,
field-to-field and cross-dataset references stopped being special cases —
they are the same lookup at a different depth.

The cost worth planning for is that the feature makes reference cycles
expressible, so a cycle check has to live somewhere. It is much cheaper to
decide that cycles are invalid when the feature is specified than to discover
it per-converter afterwards. The same restriction is also what would keep
datasets-on-datasets safe later: if a dataset's fields may only name its own
source columns, the layering stays acyclic by construction.

On the tie-break rule — good that it is explicit. Since a field and a source
column sharing a name is the common case rather than the edge case, it seems
worth having validate.py flag the shadowing cases rather than only defining
which one wins. The documents most likely to be wrong today are the ones
written before the rule existed, and they will not announce themselves.

For context, since this is my first comment here: I am CTO at XSOLCORP KOREA.
We build ElastiCORE, a model-driven platform whose YAML DSL is the authoring
format its persistence layer and APIs are generated from
(https://www.elasticore.io/docs/dsl-reference/entities). We have no Ossie
support shipped today; we are planning a bidirectional converter and intend to
contribute it here, which is why this PR matters to us — it is a precondition
for ours, not a detail.

Happy to review further revisions, particularly on round-trip and validation.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants