Skip to content

docs(bi-sql-examples): add ThoughtSpot SQL captures - #438

Merged
jbonofre merged 4 commits into
apache:mainfrom
djwaldo:feat/bi-sql-examples-thoughtspot
Oct 1, 2026
Merged

jbonofre merged 4 commits into
apache:mainfrom
djwaldo:feat/bi-sql-examples-thoughtspot

Conversation

@djwaldo

@djwaldo djwaldo commented Sep 21, 2026

Copy link
Copy Markdown
Contributor

Summary

Adds bi-sql-examples/thoughtspot/, a sibling to the existing tableau/ folder: real SQL
that ThoughtSpot generated while querying metrics defined outside the tool, in a
warehouse-native semantic layer.

The model is a six-table star with a Databricks Metric View over it. setup.sql declares
the tables and the view, and also carries a Snowflake Semantic View over the same star so
the two vendors' expressions can be compared. No data ships — the queries are reproduced
for their SQL shape rather than their results, so none of them depends on the original
rows.

Twenty-six queries across eight feature areas. Twenty-two are saved Answers captured with
POST /api/rest/2.0/metadata/answer/sql; four are AgentQL (Semantic SQL) statements
compiled by the same generator. No SQL in the folder was written by hand. Each
statement is verbatim apart from three disclosed edits (the generator's own provenance
comments removed, trailing whitespace stripped, a terminator added).

Both properties were checked mechanically rather than asserted: every statement is
byte-identical to its capture, and every statement executes against a metric view rebuilt
from this setup.sql text.

What the captures show

A few things in them bear on what a SQL interface to reusable semantics has to decide:

  • The same measure means different things on two layers over the same star, and nothing
    detects it.
    category_quantity is a dimension on the Databricks Metric View and a
    metric on the Snowflake Semantic View. Under the same filter, Databricks returns the
    whole category's units while Snowflake returns the filtered total — the window is
    evaluated before the filter on one and after it on the other. Neither errors; the numbers
    simply differ. Verified against both deployed layers (04-lod-two-stage-aggregation.sql).
  • A measure can be declared so that it is expressible where plain SQL would refuse it.
    A windowed measure declared as a dimension is grouped by and filtered on in WHERE,
    neither of which a window function permits at the same query level.
  • A top-N sub-select carries no sort tiebreak, so which rows survive LIMIT at a tie is
    undetermined. A folded top-N whose search asked for a sort does carry one — so the
    tiebreak comes from the request, confirmed against a sibling Answer differing only in
    that clause (03-top-n-and-subselects.sql).
  • Scalar functions are rewritten four ways — passed through, type renamed, substituted,
    or expanded inline. Two of the substitutions replace functions Databricks already has
    (zeroifnull, split_part), so they are choices rather than gaps. SELECT DISTINCT
    becomes GROUP BY every time (07-scalar-and-string-functions.sql).
  • Little of the visual form reaches the SQL. A pivot and a straight table built from
    byte-identical searches compile to the same shape, differing only in SELECT order — and
    because aliases are positional, that alone makes ca_2 a different measure in each
    (08-result-shape.sql).
  • The generator guards constants — a literal escaped through three nested replace()
    calls, NULLIF(3,0) eleven times in one statement, case 1 when 1 then 1 else … with an
    unreachable arm carrying live column references.

Three properties matter if the statements are reused rather than read: they are dated (a
relative date resolves at generation time to bare constants, with no CURRENT_DATE
anywhere); column aliases are positional, not stable; and a conditional is decided before
generation with the losing arm discarded, so the SQL cannot tell you whether a switch
worked or its condition never matched.

Related Issues

Discussion: #436 — opened before this PR per CONTRIBUTING.md. It raises
three open questions (whether a second subfolder on this pattern is wanted, whether a
runnable setup.sql is acceptable alongside the Tableau folder's illustrative one, and
whether the demo schema should be renamed to something neutral). Happy to reshape this
PR on any of them.

Checklist

Documentation

  • New behaviours are documented with examples

Examples

  • Adds to bi-sql-examples/, alongside the existing tableau/ folder

Tests

  • No CI workflow globs bi-sql-examples/, so none runs against this path. In place of
    that, every statement was executed against a metric view built from this PR's own
    setup.sql, and every statement was diffed against its capture.

Compliance

  • ASF license headers present on all new files
  • No third-party dependencies added

Not applicable: Specification, Ontology, Converters, Validation — this changes none of
core-spec/, ontology/, converters/ or validation/.

🤖 Generated with Claude Code

djwaldo and others added 4 commits September 21, 2026 18:04
Creates the model the captured queries run against: a six-table Dunder
Mifflin star and two warehouse-native semantic layers over it — a
Databricks Metric View and a Snowflake Semantic View.

Both layers are reduced to the columns these examples exercise, so every
object declared is used. They are otherwise as deployed, and they are not
the same surface: category_quantity is a dimension on Databricks and a
metric on Snowflake, and their comments describe it differently.

No data ships with it; the captures are reproduced for their SQL shape,
not their results.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Twenty-six queries across eight feature areas, none written by hand.
Twenty-two are saved Answers captured with metadata/answer/sql; four are
AgentQL statements compiled through the same generator. Every one was
executed against the model to confirm it is valid, not merely verbatim.

Each is reproduced as generated apart from removing the generator's own
provenance comments, stripping trailing whitespace, and adding a statement
terminator. The comment above each query records what the shape shows and,
where a Tableau counterpart exists, how the two differ.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Covers what the folder is, why the query shapes matter to a SQL interface
for reusable semantics, how the SQL was captured, and three properties of
the captures that matter if they are reused: they are dated, their column
aliases are positional rather than stable, and a parameter is resolved
before generation.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
@jbonofre
jbonofre merged commit d5e9ae6 into apache:main Oct 1, 2026
4 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants