Skip to content

Multi-Stage: Add UNNEST / CROSS JOIN UNNEST support - #17168

Merged
xiangfu0 merged 3 commits into
apache:masterfrom
xiangfu0:feat/support-unnest
Dec 10, 2025
Merged

xiangfu0 merged 3 commits into
apache:masterfrom
xiangfu0:feat/support-unnest

Conversation

@xiangfu0

@xiangfu0 xiangfu0 commented Nov 8, 2025 •

Copy link
Copy Markdown
Contributor

What's new

  • Introduces native support for array unnesting via CROSS JOIN UNNEST(...), with optional WITH ORDINALITY, in the multi-stage query engine.

Implementation

  • Wire UNNEST through the planner: new plan proto payload, serializer/deserializer support, visitors, fragmenter/sorter/validation updates, and Rel-to-plan conversion for both correlated and plain Uncollect shapes

  • Introduce UnnestNode/UnnestOperator plus plan-to-op-chain plumbing so runtime stages can expand array expressions (with optional ordinality)

  • Add coverage with logical planner unit tests, runtime operator tests, and a multi-stage integration test exercising CROSS JOIN UNNEST with and without ordinality

SQL syntax

  • Basic (single array, no ordinality)
    SELECT COUNT(*) FROM myTable CROSS JOIN UNNEST(longArrayCol) AS u(elem)

  • Projecting element values
    SELECT intCol, u.elem FROM myTable CROSS JOIN UNNEST(stringArrayCol) AS u(elem)

  • Single array with ordinality (1-based index)
    SELECT intCol, u.elem, u.idx FROM myTable CROSS JOIN UNNEST(stringArrayCol) WITH ORDINALITY AS u(elem, idx)

  • Filtering on ordinality
    SELECT COUNT(u.elem), SUM(u.idx) FROM myTable CROSS JOIN UNNEST(stringArrayCol) WITH ORDINALITY AS u(elem, idx) WHERE idx = 2

  • Multiple arrays (zip alignment, no ordinality)
    SELECT intCol, u.longValue, u.stringValue FROM myTable CROSS JOIN UNNEST(longArrayCol, stringArrayCol) AS u(longValue, stringValue)

  • Multiple arrays with ordinality
    SELECT COUNT(u.longValue), SUM(u.ord) FROM myTable CROSS JOIN UNNEST(longArrayCol, stringArrayCol) WITH ORDINALITY AS u(longValue, stringValue, ord) WHERE ord = 3

Reference

https://www.postgresql.org/docs/current/queries-table-expressions.html#QUERIES-TABLEFUNCTIONS

Sample Queries

  • Basic UNNEST (CROSS JOIN)
SELECT COUNT(*)
FROM myTable
CROSS JOIN UNNEST(longArrayCol) AS u(elem);
  • Select base column + element
SELECT intCol, u.elem
FROM myTable
CROSS JOIN UNNEST(stringArrayCol) AS u(elem)
ORDER BY intCol;
  • WITH ORDINALITY (1-based index), ordered
SELECT intCol, u.elem, u.idx
FROM myTable
CROSS JOIN UNNEST(stringArrayCol) WITH ORDINALITY AS u(elem, idx)
ORDER BY intCol, u.idx;
  • Filter on ordinality
SELECT COUNT(u.elem), SUM(u.idx)
FROM myTable
CROSS JOIN UNNEST(stringArrayCol) WITH ORDINALITY AS u(elem, idx)
WHERE idx = 2;
  • Filter on element value
SELECT intCol
FROM myTable
CROSS JOIN UNNEST(longArrayCol) AS u(val)
WHERE u.val > 10;
  • Alias variation for WITH ORDINALITY
SELECT SUM(w.ord)
FROM myTable
CROSS JOIN UNNEST(stringArrayCol) WITH ORDINALITY AS w(s, ord);
  • Selecting subset of outputs (outer-project shape)
-- **Only element**
SELECT u.elem
FROM myTable
CROSS JOIN UNNEST(stringArrayCol) WITH ORDINALITY AS u(elem, idx);

-- **Only ordinality**
SELECT u.idx
FROM myTable
CROSS JOIN UNNEST(stringArrayCol) WITH ORDINALITY AS u(elem, idx);
  • Group by base column and element
SELECT groupKey, u.elem, COUNT(*)
FROM myTable
CROSS JOIN UNNEST(stringArrayCol) AS u(elem)
GROUP BY groupKey, u.elem;
  • Using explicit table alias for correlated reference
SELECT t.intCol, e
FROM myTable t
CROSS JOIN UNNEST(t.stringArrayCol) AS u(e);
  • Multiple arrays (zip alignment without ordinality)

    SELECT intCol, u.longValue, u.stringValue
    FROM myTable
    CROSS JOIN UNNEST(longArrayCol, stringArrayCol) AS u(longValue, stringValue)
    ORDER BY intCol, u.longValue;
  • Multiple arrays with ordinality + filter

    SELECT COUNT(u.longValue), SUM(u.ord)
    FROM myTable
    CROSS JOIN UNNEST(longArrayCol, stringArrayCol) WITH ORDINALITY AS u(longValue, stringValue, ord)
    WHERE ord = 3;
  • Partial projection with multi-array ordinality

    SELECT u.stringValue, u.ord
    FROM myTable
    CROSS JOIN UNNEST(longArrayCol, stringArrayCol) WITH ORDINALITY AS u(longValue, stringValue, ord);
  • Aggregate on ordinality only

    SELECT SUM(u.ord)
    FROM myTable
    CROSS JOIN UNNEST(stringArrayCol) WITH ORDINALITY AS u(elem, ord);
  • Outer WHERE using base column plus element

    SELECT intCol, u.elem
    FROM myTable
    CROSS JOIN UNNEST(stringArrayCol) AS u(elem)
    WHERE intCol BETWEEN 5 AND 10 AND u.elem <> 'x';
  • Group aggregation over zipped arrays

    SELECT groupKey, COUNT(*), SUM(u.longValue)
    FROM myTable
    CROSS JOIN UNNEST(longArrayCol, stringArrayCol) AS u(longValue, stringValue)
    GROUP BY groupKey;

@xiangfu0
xiangfu0 requested a review from Copilot November 8, 2025 22:28
@xiangfu0 xiangfu0 added the multi-stage Related to the multi-stage query engine label Nov 8, 2025
@xiangfu0
xiangfu0 marked this pull request as draft November 8, 2025 22:28

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull Request Overview

This PR adds UNNEST functionality to Apache Pinot's multi-stage query engine, enabling SQL users to expand array/multi-value columns into multiple rows using standard ANSI CROSS JOIN UNNEST(expr) AS alias(column) syntax.

Key Changes

  • Introduced a new UnnestNode plan node with full serialization support
  • Implemented UnnestOperator runtime component to expand arrays/lists into multiple rows
  • Extended Calcite relational node converter to handle Uncollect and LogicalCorrelate patterns for UNNEST operations

Reviewed Changes

Copilot reviewed 24 out of 24 changed files in this pull request and generated 5 comments.

Show a summary per file
File Description
pinot-common/src/main/proto/plan.proto Added UnnestNode protobuf message definition for SerDe
pinot-query-planner/src/main/java/org/apache/pinot/query/planner/plannode/UnnestNode.java New plan node class modeling UNNEST semantics with array expression and column alias
pinot-query-planner/src/main/java/org/apache/pinot/query/planner/plannode/PlanNodeVisitor.java Added visitUnnest method to visitor interface
pinot-query-planner/src/main/java/org/apache/pinot/query/planner/serde/PlanNodeSerializer.java Serialization logic for UnnestNode to protobuf
pinot-query-planner/src/main/java/org/apache/pinot/query/planner/serde/PlanNodeDeserializer.java Deserialization logic from protobuf to UnnestNode
pinot-query-planner/src/main/java/org/apache/pinot/query/planner/logical/RelToPlanNodeConverter.java Converts Calcite Uncollect/LogicalCorrelate to UnnestNode with input ref resolution
pinot-query-planner/src/main/java/org/apache/pinot/query/planner/logical/PlanFragmenter.java Handles UnnestNode during plan fragmentation
pinot-query-planner/src/main/java/org/apache/pinot/query/planner/validation/ArrayToMvValidationVisitor.java Added validation support for UnnestNode
pinot-query-runtime/src/main/java/org/apache/pinot/query/runtime/operator/UnnestOperator.java Runtime operator that expands arrays/lists into multiple output rows
pinot-query-runtime/src/main/java/org/apache/pinot/query/runtime/operator/MultiStageOperator.java Registered UNNEST operator type with stats merging
pinot-query-runtime/src/main/java/org/apache/pinot/query/runtime/plan/PlanNodeToOpChain.java Maps UnnestNode to UnnestOperator during execution plan construction
pinot-query-runtime/src/test/java/org/apache/pinot/query/runtime/operator/UnnestOperatorTest.java Unit tests for UnnestOperator covering primitive arrays and lists
pinot-integration-tests/src/test/java/org/apache/pinot/integration/tests/custom/UnnestIntegrationTest.java Integration tests validating COUNT and SELECT with CROSS JOIN UNNEST

@xiangfu0
xiangfu0 force-pushed the feat/support-unnest branch from 49d6aca to c273b55 Compare November 8, 2025 22:49
@codecov-commenter

codecov-commenter commented Nov 8, 2025 •

Copy link
Copy Markdown

Codecov Report

❌ Patch coverage is 51.53257% with 253 lines in your changes missing coverage. Please review.
✅ Project coverage is 63.24%. Comparing base (57364b1) to head (ef0da08).

Files with missing lines Patch % Lines
.../query/planner/logical/RelToPlanNodeConverter.java 60.35% 66 Missing and 24 partials ⚠️
...e/pinot/query/runtime/operator/UnnestOperator.java 59.57% 41 Missing and 16 partials ⚠️
.../query/planner/logical/PlanNodeToRelConverter.java 0.00% 21 Missing ⚠️
...pache/pinot/query/planner/plannode/UnnestNode.java 68.42% 6 Missing and 12 partials ⚠️
...inot/query/planner/serde/PlanNodeDeserializer.java 0.00% 16 Missing ⚠️
...he/pinot/query/planner/explain/PlanNodeMerger.java 0.00% 15 Missing ⚠️
.../query/planner/logical/EquivalentStagesFinder.java 0.00% 8 Missing ⚠️
.../pinot/query/planner/serde/PlanNodeSerializer.java 0.00% 8 Missing ⚠️
...he/pinot/query/runtime/plan/PlanNodeToOpChain.java 0.00% 6 Missing ⚠️
.../pinot/query/planner/plannode/PlanNodeVisitor.java 0.00% 3 Missing ⚠️
... and 7 more
Additional details and impacted files
@@             Coverage Diff              @@
##             master   #17168      +/-   ##
============================================
- Coverage     63.28%   63.24%   -0.04%     
- Complexity     1433     1434       +1     
============================================
  Files          3134     3136       +2     
  Lines        186304   186825     +521     
  Branches      28457    28593     +136     
============================================
+ Hits         117898   118160     +262     
- Misses        59312    59530     +218     
- Partials       9094     9135      +41     
Flag Coverage Δ
custom-integration1 100.00% <ø> (ø)
integration 100.00% <ø> (ø)
integration1 100.00% <ø> (ø)
integration2 0.00% <ø> (ø)
java-11 63.19% <51.53%> (-0.03%) ⬇️
java-21 63.20% <51.53%> (-0.05%) ⬇️
temurin 63.24% <51.53%> (-0.04%) ⬇️
unittests 63.24% <51.53%> (-0.04%) ⬇️
unittests1 55.68% <51.53%> (-0.01%) ⬇️
unittests2 33.79% <0.00%> (-0.12%) ⬇️

Flags with carried forward coverage won't be shown. Click here to find out more.

☔ View full report in Codecov by Sentry.
📢 Have feedback on the report? Share it here.

🚀 New features to boost your workflow:
  • 📦 JS Bundle Analysis: Save yourself from yourself by tracking and limiting bundle sizes in JS merges.

@xiangfu0
xiangfu0 force-pushed the feat/support-unnest branch from c273b55 to 0ab7467 Compare November 9, 2025 09:49
@xiangfu0
xiangfu0 marked this pull request as ready for review November 9, 2025 10:02
@xiangfu0
xiangfu0 requested a review from Copilot November 9, 2025 19:49

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull Request Overview

Copilot reviewed 25 out of 25 changed files in this pull request and generated 1 comment.

@xiangfu0
xiangfu0 force-pushed the feat/support-unnest branch 2 times, most recently from fa6f9cc to bb713cb Compare November 18, 2025 20:24
@xiangfu0
xiangfu0 requested a review from Copilot November 18, 2025 20:24

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull Request Overview

Copilot reviewed 25 out of 25 changed files in this pull request and generated 3 comments.

Comments suppressed due to low confidence (1)

pinot-query-planner/src/main/java/org/apache/pinot/query/planner/logical/RelToPlanNodeConverter.java:1

  • Corrected spacing by removing extra semicolon on next line.
/**

@xiangfu0
xiangfu0 force-pushed the feat/support-unnest branch from bb713cb to 89fa166 Compare November 24, 2025 20:48
@xiangfu0
xiangfu0 requested a review from Copilot November 24, 2025 21:24

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull request overview

Copilot reviewed 25 out of 25 changed files in this pull request and generated 3 comments.

@xiangfu0
xiangfu0 force-pushed the feat/support-unnest branch from 89fa166 to 497c9dd Compare November 24, 2025 21:36
@xiangfu0
xiangfu0 requested review from Jackie-Jiang and Copilot and removed request for Copilot November 24, 2025 21:36
Comment thread pinot-common/src/main/proto/plan.proto Outdated
Comment thread pinot-common/src/main/proto/plan.proto Outdated
@xiangfu0
xiangfu0 force-pushed the feat/support-unnest branch 2 times, most recently from 8adfd69 to 5d2ecd2 Compare November 26, 2025 13:21

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull request overview

Copilot reviewed 27 out of 27 changed files in this pull request and generated 6 comments.

@xiangfu0
xiangfu0 force-pushed the feat/support-unnest branch from 5d2ecd2 to 09adedb Compare November 26, 2025 13:52
@xiangfu0
xiangfu0 requested a review from Copilot November 26, 2025 22:10
@xiangfu0
xiangfu0 force-pushed the feat/support-unnest branch from c750a28 to ef0da08 Compare December 6, 2025 19:41
@xiangfu0
xiangfu0 merged commit 0d787b9 into apache:master Dec 10, 2025
35 of 36 checks passed
@xiangfu0
xiangfu0 deleted the feat/support-unnest branch December 10, 2025 01:28
@xiangfu0 xiangfu0 added the join Related to JOIN operations label Mar 20, 2026
xiangfu0 pushed a commit to Akanksha-kedia/pinot-docs that referenced this pull request May 13, 2026
Document the Unnest operator introduced in Pinot 1.5.0 (apache/pinot#17168).
Covers basic unnesting, WITH ORDINALITY, multiple arrays, filtering and
aggregating after unnest, and explain plan output.

Also replace the long-standing _Feature TBD_ placeholder in
ingestion-level-transformations.md with a pointer to the new page.
xiangfu0 pushed a commit to pinot-contrib/pinot-docs that referenced this pull request May 13, 2026
## Description

Adds a complete operator reference page for the Unnest operator (`CROSS
JOIN UNNEST` syntax) in the multi-stage query engine, and replaces a
long-standing `_Feature TBD_` placeholder in the ingestion docs with a
concrete example linking to the new page.

## Related Issue

The CROSS JOIN UNNEST operator was introduced in Pinot 1.5.0
(apache/pinot#17168). The operator-types section had no documentation
page for it despite the 1.5.0 release notes listing it as a major MSE
feature. Additionally, `ingestion-level-transformations.md` contained an
explicit "_Feature TBD_" stub for the "Extract attributes from complex
objects" use case, which CROSS JOIN UNNEST now addresses.

## Changes Made

- **New file**:
`build-with-pinot/querying-and-sql/multi-stage-query/operator-types/unnest.md`
— full operator page covering:
  - Basic unnesting syntax
  - WITH ORDINALITY for positional indexing
  - Unnesting multiple arrays
  - Filtering on unnested elements
  - Aggregating after unnest
  - NULL/empty array handling semantics
  - Implementation details (streaming nature)
  - Stats (executionTimeMs, emittedRows)
  - Explain plan attributes (LogicalCorrelate)
- **SUMMARY.md**: Added Unnest entry in alphabetical order (between
Union and Window)
- **operator-types/README.md**: Added Unnest to the operator list
- **ingestion-level-transformations.md**: Replaced `_Feature TBD_`
placeholder in "Extract attributes from complex objects" with a CROSS
JOIN UNNEST example and link to the new page

## Testing Done

- [x] Verified SUMMARY.md entry renders correctly in local GitBook
preview
- [x] Verified cross-page links resolve (relative paths from ingestion
page to operator page)
- [x] Confirmed SQL examples execute successfully against a Pinot 1.5.0
cluster with MSE enabled
- [x] Checked alphabetical ordering is correct in both SUMMARY.md and
README.md

## Checklist

- [x] Follows existing operator page structure (description, syntax,
hints, stats, explain attributes)
- [x] No broken links introduced
- [x] New page has GitBook frontmatter (`description` field)
- [x] Hint block clearly notes MSE-only requirement
cypherean pushed a commit to cypherean/pinot that referenced this pull request Jun 24, 2026
* Support unnest operator

* address comments

* remove alias from the unnest proto
xiangfu0 added a commit to xiangfu0/pinot that referenced this pull request Aug 20, 2026
Adds row pattern recognition to the multi-stage query engine, following the
same shape as the UNNEST support added in apache#17168: a new plan node, an
exchange-insertion rule, and an intermediate-stage operator.

Pinot's parser already accepted the full MATCH_RECOGNIZE grammar (Parser.jj
carries Calcite's Babel production). What was missing was operator-table
registration, validation, a plan node, and all of execution.

Front end
- Register PREV/NEXT/FIRST/LAST/CLASSIFIER/MATCH_NUMBER/RUNNING/FINAL in
  PinotOperatorTable, which is a strict allow-list.
- New MatchRecognizeValidator runs on the SqlNode tree before conversion. It
  rewrites an OMITTED AFTER MATCH clause to SKIP PAST LAST ROW: Calcite
  substitutes SKIP TO NEXT ROW, but SQL:2016, Trino, Snowflake and Oracle all
  default to SKIP PAST LAST ROW. The two differ in whether matches overlap, so
  a query ported from another engine would otherwise silently return different
  rows. Conversion erases the omitted-vs-explicit distinction, so the rewrite
  has to happen here.
- Deferred constructs are rejected at planning time with actionable messages
  rather than silently mis-executing: ALL ROWS PER MATCH, SUBSET, PERMUTE,
  pattern exclusions, WITHIN, aggregates in DEFINE, NULLS FIRST/LAST, and
  ORDER BY / PARTITION BY on expressions (the last of which otherwise fail
  inside SqlToRelConverter with AssertionError or ClassCastException).

Plan and wire format
- MatchNode at plan.proto tag 19, encoding the pattern as a self-contained
  recursive RowPattern with a symbol table rather than a RexCall tree, with
  field numbers reserved for SUBSET and WITHIN.
- PatternFieldRef in expressions.proto, so RexPatternFieldRef can no longer
  degrade to a plain InputRef and produce wrong-but-type-correct results.
- PinotMatchExchangeNodeInsertRule hash-distributes on the PARTITION BY keys
  and prepends them to the sort collation, so rows arrive clustered and the
  operator can match and flush one partition at a time. A missing PARTITION BY
  is rejected by default, since it collapses the table onto one worker.

Runtime
- PatternToNfaCompiler builds an NFA with prioritized transitions (alternation
  in source order; greedy takes the loop edge first, reluctant the exit edge;
  {n,m} via counter registers rather than state unrolling), so depth-first
  traversal yields the SQL:2016 preferred match first.
- MatchOperator evaluates DEFINE predicates over a classifier tape supporting
  PREV/NEXT/FIRST/LAST/CLASSIFIER/MATCH_NUMBER, emits MEASURES, and advances
  per the AFTER MATCH SKIP mode.
- Guardrails throw rather than truncate, since a truncated pattern result is a
  silently wrong one: maxRowsInMatch, maxStepsPerMatchAttempt, and an
  empty-cycle guard.

v1 covers PARTITION BY, mandatory ORDER BY, MEASURES, ONE ROW PER MATCH, all
four AFTER MATCH SKIP modes, the full pattern algebra including reluctant
quantifiers and anchors, and single-variable aggregates in MEASURES.

MATCH_RECOGNIZE is not yet supported under the v2 physical optimizer or lite
mode; queries there are covered by ignore flags rather than silently wrong
results.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
xiangfu0 added a commit to xiangfu0/pinot that referenced this pull request Aug 24, 2026
Adds row pattern recognition to the multi-stage query engine, following the
same shape as the UNNEST support added in apache#17168: a new plan node, an
exchange-insertion rule, and an intermediate-stage operator.

Pinot's parser already accepted the full MATCH_RECOGNIZE grammar (Parser.jj
carries Calcite's Babel production). What was missing was operator-table
registration, validation, a plan node, and all of execution.

Front end
- Register PREV/NEXT/FIRST/LAST/CLASSIFIER/MATCH_NUMBER/RUNNING/FINAL in
  PinotOperatorTable, which is a strict allow-list.
- New MatchRecognizeValidator runs on the SqlNode tree before conversion. It
  rewrites an OMITTED AFTER MATCH clause to SKIP PAST LAST ROW: Calcite
  substitutes SKIP TO NEXT ROW, but SQL:2016, Trino, Snowflake and Oracle all
  default to SKIP PAST LAST ROW. The two differ in whether matches overlap, so
  a query ported from another engine would otherwise silently return different
  rows. Conversion erases the omitted-vs-explicit distinction, so the rewrite
  has to happen here.
- Deferred constructs are rejected at planning time with actionable messages
  rather than silently mis-executing: ALL ROWS PER MATCH, SUBSET, PERMUTE,
  pattern exclusions, WITHIN, aggregates in DEFINE, NULLS FIRST/LAST, and
  ORDER BY / PARTITION BY on expressions (the last of which otherwise fail
  inside SqlToRelConverter with AssertionError or ClassCastException).

Plan and wire format
- MatchNode at plan.proto tag 19, encoding the pattern as a self-contained
  recursive RowPattern with a symbol table rather than a RexCall tree, with
  field numbers reserved for SUBSET and WITHIN.
- PatternFieldRef in expressions.proto, so RexPatternFieldRef can no longer
  degrade to a plain InputRef and produce wrong-but-type-correct results.
- PinotMatchExchangeNodeInsertRule hash-distributes on the PARTITION BY keys
  and prepends them to the sort collation, so rows arrive clustered and the
  operator can match and flush one partition at a time. A missing PARTITION BY
  is rejected by default, since it collapses the table onto one worker.

Runtime
- PatternToNfaCompiler builds an NFA with prioritized transitions (alternation
  in source order; greedy takes the loop edge first, reluctant the exit edge;
  {n,m} via counter registers rather than state unrolling), so depth-first
  traversal yields the SQL:2016 preferred match first.
- MatchOperator evaluates DEFINE predicates over a classifier tape supporting
  PREV/NEXT/FIRST/LAST/CLASSIFIER/MATCH_NUMBER, emits MEASURES, and advances
  per the AFTER MATCH SKIP mode.
- Guardrails throw rather than truncate, since a truncated pattern result is a
  silently wrong one: maxRowsInMatch, maxStepsPerMatchAttempt, and an
  empty-cycle guard.

v1 covers PARTITION BY, mandatory ORDER BY, MEASURES, ONE ROW PER MATCH, all
four AFTER MATCH SKIP modes, the full pattern algebra including reluctant
quantifiers and anchors, and single-variable aggregates in MEASURES.

MATCH_RECOGNIZE is not yet supported under the v2 physical optimizer or lite
mode; queries there are covered by ignore flags rather than silently wrong
results.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
xiangfu0 added a commit to xiangfu0/pinot that referenced this pull request Aug 25, 2026
Adds row pattern recognition to the multi-stage query engine, following the
same shape as the UNNEST support added in apache#17168: a new plan node, an
exchange-insertion rule, and an intermediate-stage operator.

Pinot's parser already accepted the full MATCH_RECOGNIZE grammar (Parser.jj
carries Calcite's Babel production). What was missing was operator-table
registration, validation, a plan node, and all of execution.

Front end
- Register PREV/NEXT/FIRST/LAST/CLASSIFIER/MATCH_NUMBER/RUNNING/FINAL in
  PinotOperatorTable, which is a strict allow-list.
- New MatchRecognizeValidator runs on the SqlNode tree before conversion. It
  rewrites an OMITTED AFTER MATCH clause to SKIP PAST LAST ROW: Calcite
  substitutes SKIP TO NEXT ROW, but SQL:2016, Trino, Snowflake and Oracle all
  default to SKIP PAST LAST ROW. The two differ in whether matches overlap, so
  a query ported from another engine would otherwise silently return different
  rows. Conversion erases the omitted-vs-explicit distinction, so the rewrite
  has to happen here.
- Deferred constructs are rejected at planning time with actionable messages
  rather than silently mis-executing: ALL ROWS PER MATCH, SUBSET, PERMUTE,
  pattern exclusions, WITHIN, aggregates in DEFINE, NULLS FIRST/LAST, and
  ORDER BY / PARTITION BY on expressions (the last of which otherwise fail
  inside SqlToRelConverter with AssertionError or ClassCastException).

Plan and wire format
- MatchNode at plan.proto tag 19, encoding the pattern as a self-contained
  recursive RowPattern with a symbol table rather than a RexCall tree, with
  field numbers reserved for SUBSET and WITHIN.
- PatternFieldRef in expressions.proto, so RexPatternFieldRef can no longer
  degrade to a plain InputRef and produce wrong-but-type-correct results.
- PinotMatchExchangeNodeInsertRule hash-distributes on the PARTITION BY keys
  and prepends them to the sort collation, so rows arrive clustered and the
  operator can match and flush one partition at a time. A missing PARTITION BY
  is rejected by default, since it collapses the table onto one worker.

Runtime
- PatternToNfaCompiler builds an NFA with prioritized transitions (alternation
  in source order; greedy takes the loop edge first, reluctant the exit edge;
  {n,m} via counter registers rather than state unrolling), so depth-first
  traversal yields the SQL:2016 preferred match first.
- MatchOperator evaluates DEFINE predicates over a classifier tape supporting
  PREV/NEXT/FIRST/LAST/CLASSIFIER/MATCH_NUMBER, emits MEASURES, and advances
  per the AFTER MATCH SKIP mode.
- Guardrails throw rather than truncate, since a truncated pattern result is a
  silently wrong one: maxRowsInMatch, maxStepsPerMatchAttempt, and an
  empty-cycle guard.

v1 covers PARTITION BY, mandatory ORDER BY, MEASURES, ONE ROW PER MATCH, all
four AFTER MATCH SKIP modes, the full pattern algebra including reluctant
quantifiers and anchors, and single-variable aggregates in MEASURES.

MATCH_RECOGNIZE is not yet supported under the v2 physical optimizer or lite
mode; queries there are covered by ignore flags rather than silently wrong
results.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
xiangfu0 added a commit to xiangfu0/pinot that referenced this pull request Aug 26, 2026
Adds row pattern recognition to the multi-stage query engine, following the
same shape as the UNNEST support added in apache#17168: a new plan node, an
exchange-insertion rule, and an intermediate-stage operator.

Pinot's parser already accepted the full MATCH_RECOGNIZE grammar (Parser.jj
carries Calcite's Babel production). What was missing was operator-table
registration, validation, a plan node, and all of execution.

Front end
- Register PREV/NEXT/FIRST/LAST/CLASSIFIER/MATCH_NUMBER/RUNNING/FINAL in
  PinotOperatorTable, which is a strict allow-list.
- New MatchRecognizeValidator runs on the SqlNode tree before conversion. It
  rewrites an OMITTED AFTER MATCH clause to SKIP PAST LAST ROW: Calcite
  substitutes SKIP TO NEXT ROW, but SQL:2016, Trino, Snowflake and Oracle all
  default to SKIP PAST LAST ROW. The two differ in whether matches overlap, so
  a query ported from another engine would otherwise silently return different
  rows. Conversion erases the omitted-vs-explicit distinction, so the rewrite
  has to happen here.
- Deferred constructs are rejected at planning time with actionable messages
  rather than silently mis-executing: ALL ROWS PER MATCH, SUBSET, PERMUTE,
  pattern exclusions, WITHIN, aggregates in DEFINE, NULLS FIRST/LAST, and
  ORDER BY / PARTITION BY on expressions (the last of which otherwise fail
  inside SqlToRelConverter with AssertionError or ClassCastException).

Plan and wire format
- MatchNode at plan.proto tag 19, encoding the pattern as a self-contained
  recursive RowPattern with a symbol table rather than a RexCall tree, with
  field numbers reserved for SUBSET and WITHIN.
- PatternFieldRef in expressions.proto, so RexPatternFieldRef can no longer
  degrade to a plain InputRef and produce wrong-but-type-correct results.
- PinotMatchExchangeNodeInsertRule hash-distributes on the PARTITION BY keys
  and prepends them to the sort collation, so rows arrive clustered and the
  operator can match and flush one partition at a time. A missing PARTITION BY
  is rejected by default, since it collapses the table onto one worker.

Runtime
- PatternToNfaCompiler builds an NFA with prioritized transitions (alternation
  in source order; greedy takes the loop edge first, reluctant the exit edge;
  {n,m} via counter registers rather than state unrolling), so depth-first
  traversal yields the SQL:2016 preferred match first.
- MatchOperator evaluates DEFINE predicates over a classifier tape supporting
  PREV/NEXT/FIRST/LAST/CLASSIFIER/MATCH_NUMBER, emits MEASURES, and advances
  per the AFTER MATCH SKIP mode.
- Guardrails throw rather than truncate, since a truncated pattern result is a
  silently wrong one: maxRowsInMatch, maxStepsPerMatchAttempt, and an
  empty-cycle guard.

v1 covers PARTITION BY, mandatory ORDER BY, MEASURES, ONE ROW PER MATCH, all
four AFTER MATCH SKIP modes, the full pattern algebra including reluctant
quantifiers and anchors, and single-variable aggregates in MEASURES.

MATCH_RECOGNIZE is not yet supported under the v2 physical optimizer or lite
mode; queries there are covered by ignore flags rather than silently wrong
results.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
xiangfu0 added a commit to xiangfu0/pinot that referenced this pull request Aug 27, 2026
Adds row pattern recognition to the multi-stage query engine, following the
same shape as the UNNEST support added in apache#17168: a new plan node, an
exchange-insertion rule, and an intermediate-stage operator.

Pinot's parser already accepted the full MATCH_RECOGNIZE grammar (Parser.jj
carries Calcite's Babel production). What was missing was operator-table
registration, validation, a plan node, and all of execution.

Front end
- Register PREV/NEXT/FIRST/LAST/CLASSIFIER/MATCH_NUMBER/RUNNING/FINAL in
  PinotOperatorTable, which is a strict allow-list.
- New MatchRecognizeValidator runs on the SqlNode tree before conversion. It
  rewrites an OMITTED AFTER MATCH clause to SKIP PAST LAST ROW: Calcite
  substitutes SKIP TO NEXT ROW, but SQL:2016, Trino, Snowflake and Oracle all
  default to SKIP PAST LAST ROW. The two differ in whether matches overlap, so
  a query ported from another engine would otherwise silently return different
  rows. Conversion erases the omitted-vs-explicit distinction, so the rewrite
  has to happen here.
- Deferred constructs are rejected at planning time with actionable messages
  rather than silently mis-executing: ALL ROWS PER MATCH, SUBSET, PERMUTE,
  pattern exclusions, WITHIN, aggregates in DEFINE, NULLS FIRST/LAST, and
  ORDER BY / PARTITION BY on expressions (the last of which otherwise fail
  inside SqlToRelConverter with AssertionError or ClassCastException).

Plan and wire format
- MatchNode at plan.proto tag 19, encoding the pattern as a self-contained
  recursive RowPattern with a symbol table rather than a RexCall tree, with
  field numbers reserved for SUBSET and WITHIN.
- PatternFieldRef in expressions.proto, so RexPatternFieldRef can no longer
  degrade to a plain InputRef and produce wrong-but-type-correct results.
- PinotMatchExchangeNodeInsertRule hash-distributes on the PARTITION BY keys
  and prepends them to the sort collation, so rows arrive clustered and the
  operator can match and flush one partition at a time. A missing PARTITION BY
  is rejected by default, since it collapses the table onto one worker.

Runtime
- PatternToNfaCompiler builds an NFA with prioritized transitions (alternation
  in source order; greedy takes the loop edge first, reluctant the exit edge;
  {n,m} via counter registers rather than state unrolling), so depth-first
  traversal yields the SQL:2016 preferred match first.
- MatchOperator evaluates DEFINE predicates over a classifier tape supporting
  PREV/NEXT/FIRST/LAST/CLASSIFIER/MATCH_NUMBER, emits MEASURES, and advances
  per the AFTER MATCH SKIP mode.
- Guardrails throw rather than truncate, since a truncated pattern result is a
  silently wrong one: maxRowsInMatch, maxStepsPerMatchAttempt, and an
  empty-cycle guard.

v1 covers PARTITION BY, mandatory ORDER BY, MEASURES, ONE ROW PER MATCH, all
four AFTER MATCH SKIP modes, the full pattern algebra including reluctant
quantifiers and anchors, and single-variable aggregates in MEASURES.

MATCH_RECOGNIZE is not yet supported under the v2 physical optimizer or lite
mode; queries there are covered by ignore flags rather than silently wrong
results.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
xiangfu0 added a commit to xiangfu0/pinot that referenced this pull request Aug 27, 2026
Adds row pattern recognition to the multi-stage query engine, following the
same shape as the UNNEST support added in apache#17168: a new plan node, an
exchange-insertion rule, and an intermediate-stage operator.

Pinot's parser already accepted the full MATCH_RECOGNIZE grammar (Parser.jj
carries Calcite's Babel production). What was missing was operator-table
registration, validation, a plan node, and all of execution.

Front end
- Register PREV/NEXT/FIRST/LAST/CLASSIFIER/MATCH_NUMBER/RUNNING/FINAL in
  PinotOperatorTable, which is a strict allow-list.
- New MatchRecognizeValidator runs on the SqlNode tree before conversion. It
  rewrites an OMITTED AFTER MATCH clause to SKIP PAST LAST ROW: Calcite
  substitutes SKIP TO NEXT ROW, but SQL:2016, Trino, Snowflake and Oracle all
  default to SKIP PAST LAST ROW. The two differ in whether matches overlap, so
  a query ported from another engine would otherwise silently return different
  rows. Conversion erases the omitted-vs-explicit distinction, so the rewrite
  has to happen here.
- Deferred constructs are rejected at planning time with actionable messages
  rather than silently mis-executing: ALL ROWS PER MATCH, SUBSET, PERMUTE,
  pattern exclusions, WITHIN, aggregates in DEFINE, NULLS FIRST/LAST, and
  ORDER BY / PARTITION BY on expressions (the last of which otherwise fail
  inside SqlToRelConverter with AssertionError or ClassCastException).

Plan and wire format
- MatchNode at plan.proto tag 19, encoding the pattern as a self-contained
  recursive RowPattern with a symbol table rather than a RexCall tree, with
  field numbers reserved for SUBSET and WITHIN.
- PatternFieldRef in expressions.proto, so RexPatternFieldRef can no longer
  degrade to a plain InputRef and produce wrong-but-type-correct results.
- PinotMatchExchangeNodeInsertRule hash-distributes on the PARTITION BY keys
  and prepends them to the sort collation, so rows arrive clustered and the
  operator can match and flush one partition at a time. A missing PARTITION BY
  is rejected by default, since it collapses the table onto one worker.

Runtime
- PatternToNfaCompiler builds an NFA with prioritized transitions (alternation
  in source order; greedy takes the loop edge first, reluctant the exit edge;
  {n,m} via counter registers rather than state unrolling), so depth-first
  traversal yields the SQL:2016 preferred match first.
- MatchOperator evaluates DEFINE predicates over a classifier tape supporting
  PREV/NEXT/FIRST/LAST/CLASSIFIER/MATCH_NUMBER, emits MEASURES, and advances
  per the AFTER MATCH SKIP mode.
- Guardrails throw rather than truncate, since a truncated pattern result is a
  silently wrong one: maxRowsInMatch, maxStepsPerMatchAttempt, and an
  empty-cycle guard.

v1 covers PARTITION BY, mandatory ORDER BY, MEASURES, ONE ROW PER MATCH, all
four AFTER MATCH SKIP modes, the full pattern algebra including reluctant
quantifiers and anchors, and single-variable aggregates in MEASURES.

MATCH_RECOGNIZE is not yet supported under the v2 physical optimizer or lite
mode; queries there are covered by ignore flags rather than silently wrong
results.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
xiangfu0 added a commit to xiangfu0/pinot that referenced this pull request Aug 28, 2026
Adds row pattern recognition to the multi-stage query engine, following the
same shape as the UNNEST support added in apache#17168: a new plan node, an
exchange-insertion rule, and an intermediate-stage operator.

Pinot's parser already accepted the full MATCH_RECOGNIZE grammar (Parser.jj
carries Calcite's Babel production). What was missing was operator-table
registration, validation, a plan node, and all of execution.

Front end
- Register PREV/NEXT/FIRST/LAST/CLASSIFIER/MATCH_NUMBER/RUNNING/FINAL in
  PinotOperatorTable, which is a strict allow-list.
- New MatchRecognizeValidator runs on the SqlNode tree before conversion. It
  rewrites an OMITTED AFTER MATCH clause to SKIP PAST LAST ROW: Calcite
  substitutes SKIP TO NEXT ROW, but SQL:2016, Trino, Snowflake and Oracle all
  default to SKIP PAST LAST ROW. The two differ in whether matches overlap, so
  a query ported from another engine would otherwise silently return different
  rows. Conversion erases the omitted-vs-explicit distinction, so the rewrite
  has to happen here.
- Deferred constructs are rejected at planning time with actionable messages
  rather than silently mis-executing: ALL ROWS PER MATCH, SUBSET, PERMUTE,
  pattern exclusions, WITHIN, aggregates in DEFINE, NULLS FIRST/LAST, and
  ORDER BY / PARTITION BY on expressions (the last of which otherwise fail
  inside SqlToRelConverter with AssertionError or ClassCastException).

Plan and wire format
- MatchNode at plan.proto tag 19, encoding the pattern as a self-contained
  recursive RowPattern with a symbol table rather than a RexCall tree, with
  field numbers reserved for SUBSET and WITHIN.
- PatternFieldRef in expressions.proto, so RexPatternFieldRef can no longer
  degrade to a plain InputRef and produce wrong-but-type-correct results.
- PinotMatchExchangeNodeInsertRule hash-distributes on the PARTITION BY keys
  and prepends them to the sort collation, so rows arrive clustered and the
  operator can match and flush one partition at a time. A missing PARTITION BY
  is rejected by default, since it collapses the table onto one worker.

Runtime
- PatternToNfaCompiler builds an NFA with prioritized transitions (alternation
  in source order; greedy takes the loop edge first, reluctant the exit edge;
  {n,m} via counter registers rather than state unrolling), so depth-first
  traversal yields the SQL:2016 preferred match first.
- MatchOperator evaluates DEFINE predicates over a classifier tape supporting
  PREV/NEXT/FIRST/LAST/CLASSIFIER/MATCH_NUMBER, emits MEASURES, and advances
  per the AFTER MATCH SKIP mode.
- Guardrails throw rather than truncate, since a truncated pattern result is a
  silently wrong one: maxRowsInMatch, maxStepsPerMatchAttempt, and an
  empty-cycle guard.

v1 covers PARTITION BY, mandatory ORDER BY, MEASURES, ONE ROW PER MATCH, all
four AFTER MATCH SKIP modes, the full pattern algebra including reluctant
quantifiers and anchors, and single-variable aggregates in MEASURES.

MATCH_RECOGNIZE is not yet supported under the v2 physical optimizer or lite
mode; queries there are covered by ignore flags rather than silently wrong
results.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
xiangfu0 added a commit to xiangfu0/pinot that referenced this pull request Aug 29, 2026
Adds row pattern recognition to the multi-stage query engine, following the
same shape as the UNNEST support added in apache#17168: a new plan node, an
exchange-insertion rule, and an intermediate-stage operator.

Pinot's parser already accepted the full MATCH_RECOGNIZE grammar (Parser.jj
carries Calcite's Babel production). What was missing was operator-table
registration, validation, a plan node, and all of execution.

Front end
- Register PREV/NEXT/FIRST/LAST/CLASSIFIER/MATCH_NUMBER/RUNNING/FINAL in
  PinotOperatorTable, which is a strict allow-list.
- New MatchRecognizeValidator runs on the SqlNode tree before conversion. It
  rewrites an OMITTED AFTER MATCH clause to SKIP PAST LAST ROW: Calcite
  substitutes SKIP TO NEXT ROW, but SQL:2016, Trino, Snowflake and Oracle all
  default to SKIP PAST LAST ROW. The two differ in whether matches overlap, so
  a query ported from another engine would otherwise silently return different
  rows. Conversion erases the omitted-vs-explicit distinction, so the rewrite
  has to happen here.
- Deferred constructs are rejected at planning time with actionable messages
  rather than silently mis-executing: ALL ROWS PER MATCH, SUBSET, PERMUTE,
  pattern exclusions, WITHIN, aggregates in DEFINE, NULLS FIRST/LAST, and
  ORDER BY / PARTITION BY on expressions (the last of which otherwise fail
  inside SqlToRelConverter with AssertionError or ClassCastException).

Plan and wire format
- MatchNode at plan.proto tag 19, encoding the pattern as a self-contained
  recursive RowPattern with a symbol table rather than a RexCall tree, with
  field numbers reserved for SUBSET and WITHIN.
- PatternFieldRef in expressions.proto, so RexPatternFieldRef can no longer
  degrade to a plain InputRef and produce wrong-but-type-correct results.
- PinotMatchExchangeNodeInsertRule hash-distributes on the PARTITION BY keys
  and prepends them to the sort collation, so rows arrive clustered and the
  operator can match and flush one partition at a time. A missing PARTITION BY
  is rejected by default, since it collapses the table onto one worker.

Runtime
- PatternToNfaCompiler builds an NFA with prioritized transitions (alternation
  in source order; greedy takes the loop edge first, reluctant the exit edge;
  {n,m} via counter registers rather than state unrolling), so depth-first
  traversal yields the SQL:2016 preferred match first.
- MatchOperator evaluates DEFINE predicates over a classifier tape supporting
  PREV/NEXT/FIRST/LAST/CLASSIFIER/MATCH_NUMBER, emits MEASURES, and advances
  per the AFTER MATCH SKIP mode.
- Guardrails throw rather than truncate, since a truncated pattern result is a
  silently wrong one: maxRowsInMatch, maxStepsPerMatchAttempt, and an
  empty-cycle guard.

v1 covers PARTITION BY, mandatory ORDER BY, MEASURES, ONE ROW PER MATCH, all
four AFTER MATCH SKIP modes, the full pattern algebra including reluctant
quantifiers and anchors, and single-variable aggregates in MEASURES.

MATCH_RECOGNIZE is not yet supported under the v2 physical optimizer or lite
mode; queries there are covered by ignore flags rather than silently wrong
results.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
xiangfu0 added a commit to xiangfu0/pinot that referenced this pull request Aug 29, 2026
Adds row pattern recognition to the multi-stage query engine, following the
same shape as the UNNEST support added in apache#17168: a new plan node, an
exchange-insertion rule, and an intermediate-stage operator.

Pinot's parser already accepted the full MATCH_RECOGNIZE grammar (Parser.jj
carries Calcite's Babel production). What was missing was operator-table
registration, validation, a plan node, and all of execution.

Front end
- Register PREV/NEXT/FIRST/LAST/CLASSIFIER/MATCH_NUMBER/RUNNING/FINAL in
  PinotOperatorTable, which is a strict allow-list.
- New MatchRecognizeValidator runs on the SqlNode tree before conversion. It
  rewrites an OMITTED AFTER MATCH clause to SKIP PAST LAST ROW: Calcite
  substitutes SKIP TO NEXT ROW, but SQL:2016, Trino, Snowflake and Oracle all
  default to SKIP PAST LAST ROW. The two differ in whether matches overlap, so
  a query ported from another engine would otherwise silently return different
  rows. Conversion erases the omitted-vs-explicit distinction, so the rewrite
  has to happen here.
- Deferred constructs are rejected at planning time with actionable messages
  rather than silently mis-executing: ALL ROWS PER MATCH, SUBSET, PERMUTE,
  pattern exclusions, WITHIN, aggregates in DEFINE, NULLS FIRST/LAST, and
  ORDER BY / PARTITION BY on expressions (the last of which otherwise fail
  inside SqlToRelConverter with AssertionError or ClassCastException).

Plan and wire format
- MatchNode at plan.proto tag 19, encoding the pattern as a self-contained
  recursive RowPattern with a symbol table rather than a RexCall tree, with
  field numbers reserved for SUBSET and WITHIN.
- PatternFieldRef in expressions.proto, so RexPatternFieldRef can no longer
  degrade to a plain InputRef and produce wrong-but-type-correct results.
- PinotMatchExchangeNodeInsertRule hash-distributes on the PARTITION BY keys
  and prepends them to the sort collation, so rows arrive clustered and the
  operator can match and flush one partition at a time. A missing PARTITION BY
  is rejected by default, since it collapses the table onto one worker.

Runtime
- PatternToNfaCompiler builds an NFA with prioritized transitions (alternation
  in source order; greedy takes the loop edge first, reluctant the exit edge;
  {n,m} via counter registers rather than state unrolling), so depth-first
  traversal yields the SQL:2016 preferred match first.
- MatchOperator evaluates DEFINE predicates over a classifier tape supporting
  PREV/NEXT/FIRST/LAST/CLASSIFIER/MATCH_NUMBER, emits MEASURES, and advances
  per the AFTER MATCH SKIP mode.
- Guardrails throw rather than truncate, since a truncated pattern result is a
  silently wrong one: maxRowsInMatch, maxStepsPerMatchAttempt, and an
  empty-cycle guard.

v1 covers PARTITION BY, mandatory ORDER BY, MEASURES, ONE ROW PER MATCH, all
four AFTER MATCH SKIP modes, the full pattern algebra including reluctant
quantifiers and anchors, and single-variable aggregates in MEASURES.

MATCH_RECOGNIZE is not yet supported under the v2 physical optimizer or lite
mode; queries there are covered by ignore flags rather than silently wrong
results.
xiangfu0 added a commit to xiangfu0/pinot that referenced this pull request Aug 31, 2026
Adds row pattern recognition to the multi-stage query engine, following the
same shape as the UNNEST support added in apache#17168: a new plan node, an
exchange-insertion rule, and an intermediate-stage operator.

Pinot's parser already accepted the full MATCH_RECOGNIZE grammar (Parser.jj
carries Calcite's Babel production). What was missing was operator-table
registration, validation, a plan node, and all of execution.

Front end
- Register PREV/NEXT/FIRST/LAST/CLASSIFIER/MATCH_NUMBER/RUNNING/FINAL in
  PinotOperatorTable, which is a strict allow-list.
- New MatchRecognizeValidator runs on the SqlNode tree before conversion. It
  rewrites an OMITTED AFTER MATCH clause to SKIP PAST LAST ROW: Calcite
  substitutes SKIP TO NEXT ROW, but SQL:2016, Trino, Snowflake and Oracle all
  default to SKIP PAST LAST ROW. The two differ in whether matches overlap, so
  a query ported from another engine would otherwise silently return different
  rows. Conversion erases the omitted-vs-explicit distinction, so the rewrite
  has to happen here.
- Deferred constructs are rejected at planning time with actionable messages
  rather than silently mis-executing: ALL ROWS PER MATCH, SUBSET, PERMUTE,
  pattern exclusions, WITHIN, aggregates in DEFINE, NULLS FIRST/LAST, and
  ORDER BY / PARTITION BY on expressions (the last of which otherwise fail
  inside SqlToRelConverter with AssertionError or ClassCastException).

Plan and wire format
- MatchNode at plan.proto tag 19, encoding the pattern as a self-contained
  recursive RowPattern with a symbol table rather than a RexCall tree, with
  field numbers reserved for SUBSET and WITHIN.
- PatternFieldRef in expressions.proto, so RexPatternFieldRef can no longer
  degrade to a plain InputRef and produce wrong-but-type-correct results.
- PinotMatchExchangeNodeInsertRule hash-distributes on the PARTITION BY keys
  and prepends them to the sort collation, so rows arrive clustered and the
  operator can match and flush one partition at a time. A missing PARTITION BY
  is rejected by default, since it collapses the table onto one worker.

Runtime
- PatternToNfaCompiler builds an NFA with prioritized transitions (alternation
  in source order; greedy takes the loop edge first, reluctant the exit edge;
  {n,m} via counter registers rather than state unrolling), so depth-first
  traversal yields the SQL:2016 preferred match first.
- MatchOperator evaluates DEFINE predicates over a classifier tape supporting
  PREV/NEXT/FIRST/LAST/CLASSIFIER/MATCH_NUMBER, emits MEASURES, and advances
  per the AFTER MATCH SKIP mode.
- Guardrails throw rather than truncate, since a truncated pattern result is a
  silently wrong one: maxRowsInMatch, maxStepsPerMatchAttempt, and an
  empty-cycle guard.

v1 covers PARTITION BY, mandatory ORDER BY, MEASURES, ONE ROW PER MATCH, all
four AFTER MATCH SKIP modes, the full pattern algebra including reluctant
quantifiers and anchors, and single-variable aggregates in MEASURES.

MATCH_RECOGNIZE is not yet supported under the v2 physical optimizer or lite
mode; queries there are covered by ignore flags rather than silently wrong
results.
xiangfu0 added a commit to xiangfu0/pinot that referenced this pull request Sep 1, 2026
Adds row pattern recognition to the multi-stage query engine, following the
same shape as the UNNEST support added in apache#17168: a new plan node, an
exchange-insertion rule, and an intermediate-stage operator.

Pinot's parser already accepted the full MATCH_RECOGNIZE grammar (Parser.jj
carries Calcite's Babel production). What was missing was operator-table
registration, validation, a plan node, and all of execution.

Front end
- Register PREV/NEXT/FIRST/LAST/CLASSIFIER/MATCH_NUMBER/RUNNING/FINAL in
  PinotOperatorTable, which is a strict allow-list.
- New MatchRecognizeValidator runs on the SqlNode tree before conversion. It
  rewrites an OMITTED AFTER MATCH clause to SKIP PAST LAST ROW: Calcite
  substitutes SKIP TO NEXT ROW, but SQL:2016, Trino, Snowflake and Oracle all
  default to SKIP PAST LAST ROW. The two differ in whether matches overlap, so
  a query ported from another engine would otherwise silently return different
  rows. Conversion erases the omitted-vs-explicit distinction, so the rewrite
  has to happen here.
- Deferred constructs are rejected at planning time with actionable messages
  rather than silently mis-executing: ALL ROWS PER MATCH, SUBSET, PERMUTE,
  pattern exclusions, WITHIN, aggregates in DEFINE, NULLS FIRST/LAST, and
  ORDER BY / PARTITION BY on expressions (the last of which otherwise fail
  inside SqlToRelConverter with AssertionError or ClassCastException).

Plan and wire format
- MatchNode at plan.proto tag 19, encoding the pattern as a self-contained
  recursive RowPattern with a symbol table rather than a RexCall tree, with
  field numbers reserved for SUBSET and WITHIN.
- PatternFieldRef in expressions.proto, so RexPatternFieldRef can no longer
  degrade to a plain InputRef and produce wrong-but-type-correct results.
- PinotMatchExchangeNodeInsertRule hash-distributes on the PARTITION BY keys
  and prepends them to the sort collation, so rows arrive clustered and the
  operator can match and flush one partition at a time. A missing PARTITION BY
  is rejected by default, since it collapses the table onto one worker.

Runtime
- PatternToNfaCompiler builds an NFA with prioritized transitions (alternation
  in source order; greedy takes the loop edge first, reluctant the exit edge;
  {n,m} via counter registers rather than state unrolling), so depth-first
  traversal yields the SQL:2016 preferred match first.
- MatchOperator evaluates DEFINE predicates over a classifier tape supporting
  PREV/NEXT/FIRST/LAST/CLASSIFIER/MATCH_NUMBER, emits MEASURES, and advances
  per the AFTER MATCH SKIP mode.
- Guardrails throw rather than truncate, since a truncated pattern result is a
  silently wrong one: maxRowsInMatch, maxStepsPerMatchAttempt, and an
  empty-cycle guard.

v1 covers PARTITION BY, mandatory ORDER BY, MEASURES, ONE ROW PER MATCH, all
four AFTER MATCH SKIP modes, the full pattern algebra including reluctant
quantifiers and anchors, and single-variable aggregates in MEASURES.

MATCH_RECOGNIZE is not yet supported under the v2 physical optimizer or lite
mode; queries there are covered by ignore flags rather than silently wrong
results.
xiangfu0 added a commit to xiangfu0/pinot that referenced this pull request Sep 2, 2026
Adds row pattern recognition to the multi-stage query engine, following the
same shape as the UNNEST support added in apache#17168: a new plan node, an
exchange-insertion rule, and an intermediate-stage operator.

Pinot's parser already accepted the full MATCH_RECOGNIZE grammar (Parser.jj
carries Calcite's Babel production). What was missing was operator-table
registration, validation, a plan node, and all of execution.

Front end
- Register PREV/NEXT/FIRST/LAST/CLASSIFIER/MATCH_NUMBER/RUNNING/FINAL in
  PinotOperatorTable, which is a strict allow-list.
- New MatchRecognizeValidator runs on the SqlNode tree before conversion. It
  rewrites an OMITTED AFTER MATCH clause to SKIP PAST LAST ROW: Calcite
  substitutes SKIP TO NEXT ROW, but SQL:2016, Trino, Snowflake and Oracle all
  default to SKIP PAST LAST ROW. The two differ in whether matches overlap, so
  a query ported from another engine would otherwise silently return different
  rows. Conversion erases the omitted-vs-explicit distinction, so the rewrite
  has to happen here.
- Deferred constructs are rejected at planning time with actionable messages
  rather than silently mis-executing: ALL ROWS PER MATCH, SUBSET, PERMUTE,
  pattern exclusions, WITHIN, aggregates in DEFINE, NULLS FIRST/LAST, and
  ORDER BY / PARTITION BY on expressions (the last of which otherwise fail
  inside SqlToRelConverter with AssertionError or ClassCastException).

Plan and wire format
- MatchNode at plan.proto tag 19, encoding the pattern as a self-contained
  recursive RowPattern with a symbol table rather than a RexCall tree, with
  field numbers reserved for SUBSET and WITHIN.
- PatternFieldRef in expressions.proto, so RexPatternFieldRef can no longer
  degrade to a plain InputRef and produce wrong-but-type-correct results.
- PinotMatchExchangeNodeInsertRule hash-distributes on the PARTITION BY keys
  and prepends them to the sort collation, so rows arrive clustered and the
  operator can match and flush one partition at a time. A missing PARTITION BY
  is rejected by default, since it collapses the table onto one worker.

Runtime
- PatternToNfaCompiler builds an NFA with prioritized transitions (alternation
  in source order; greedy takes the loop edge first, reluctant the exit edge;
  {n,m} via counter registers rather than state unrolling), so depth-first
  traversal yields the SQL:2016 preferred match first.
- MatchOperator evaluates DEFINE predicates over a classifier tape supporting
  PREV/NEXT/FIRST/LAST/CLASSIFIER/MATCH_NUMBER, emits MEASURES, and advances
  per the AFTER MATCH SKIP mode.
- Guardrails throw rather than truncate, since a truncated pattern result is a
  silently wrong one: maxRowsInMatch, maxStepsPerMatchAttempt, and an
  empty-cycle guard.

v1 covers PARTITION BY, mandatory ORDER BY, MEASURES, ONE ROW PER MATCH, all
four AFTER MATCH SKIP modes, the full pattern algebra including reluctant
quantifiers and anchors, and single-variable aggregates in MEASURES.

MATCH_RECOGNIZE is not yet supported under the v2 physical optimizer or lite
mode; queries there are covered by ignore flags rather than silently wrong
results.
xiangfu0 added a commit to xiangfu0/pinot that referenced this pull request Sep 3, 2026
Adds row pattern recognition to the multi-stage query engine, following the
same shape as the UNNEST support added in apache#17168: a new plan node, an
exchange-insertion rule, and an intermediate-stage operator.

Pinot's parser already accepted the full MATCH_RECOGNIZE grammar (Parser.jj
carries Calcite's Babel production). What was missing was operator-table
registration, validation, a plan node, and all of execution.

Front end
- Register PREV/NEXT/FIRST/LAST/CLASSIFIER/MATCH_NUMBER/RUNNING/FINAL in
  PinotOperatorTable, which is a strict allow-list.
- New MatchRecognizeValidator runs on the SqlNode tree before conversion. It
  rewrites an OMITTED AFTER MATCH clause to SKIP PAST LAST ROW: Calcite
  substitutes SKIP TO NEXT ROW, but SQL:2016, Trino, Snowflake and Oracle all
  default to SKIP PAST LAST ROW. The two differ in whether matches overlap, so
  a query ported from another engine would otherwise silently return different
  rows. Conversion erases the omitted-vs-explicit distinction, so the rewrite
  has to happen here.
- Deferred constructs are rejected at planning time with actionable messages
  rather than silently mis-executing: ALL ROWS PER MATCH, SUBSET, PERMUTE,
  pattern exclusions, WITHIN, aggregates in DEFINE, NULLS FIRST/LAST, and
  ORDER BY / PARTITION BY on expressions (the last of which otherwise fail
  inside SqlToRelConverter with AssertionError or ClassCastException).

Plan and wire format
- MatchNode at plan.proto tag 19, encoding the pattern as a self-contained
  recursive RowPattern with a symbol table rather than a RexCall tree, with
  field numbers reserved for SUBSET and WITHIN.
- PatternFieldRef in expressions.proto, so RexPatternFieldRef can no longer
  degrade to a plain InputRef and produce wrong-but-type-correct results.
- PinotMatchExchangeNodeInsertRule hash-distributes on the PARTITION BY keys
  and prepends them to the sort collation, so rows arrive clustered and the
  operator can match and flush one partition at a time. A missing PARTITION BY
  is rejected by default, since it collapses the table onto one worker.

Runtime
- PatternToNfaCompiler builds an NFA with prioritized transitions (alternation
  in source order; greedy takes the loop edge first, reluctant the exit edge;
  {n,m} via counter registers rather than state unrolling), so depth-first
  traversal yields the SQL:2016 preferred match first.
- MatchOperator evaluates DEFINE predicates over a classifier tape supporting
  PREV/NEXT/FIRST/LAST/CLASSIFIER/MATCH_NUMBER, emits MEASURES, and advances
  per the AFTER MATCH SKIP mode.
- Guardrails throw rather than truncate, since a truncated pattern result is a
  silently wrong one: maxRowsInMatch, maxStepsPerMatchAttempt, and an
  empty-cycle guard.

v1 covers PARTITION BY, mandatory ORDER BY, MEASURES, ONE ROW PER MATCH, all
four AFTER MATCH SKIP modes, the full pattern algebra including reluctant
quantifiers and anchors, and single-variable aggregates in MEASURES.

MATCH_RECOGNIZE is not yet supported under the v2 physical optimizer or lite
mode; queries there are covered by ignore flags rather than silently wrong
results.
xiangfu0 added a commit to xiangfu0/pinot that referenced this pull request Sep 4, 2026
Adds row pattern recognition to the multi-stage query engine, following the
same shape as the UNNEST support added in apache#17168: a new plan node, an
exchange-insertion rule, and an intermediate-stage operator.

Pinot's parser already accepted the full MATCH_RECOGNIZE grammar (Parser.jj
carries Calcite's Babel production). What was missing was operator-table
registration, validation, a plan node, and all of execution.

Front end
- Register PREV/NEXT/FIRST/LAST/CLASSIFIER/MATCH_NUMBER/RUNNING/FINAL in
  PinotOperatorTable, which is a strict allow-list.
- New MatchRecognizeValidator runs on the SqlNode tree before conversion. It
  rewrites an OMITTED AFTER MATCH clause to SKIP PAST LAST ROW: Calcite
  substitutes SKIP TO NEXT ROW, but SQL:2016, Trino, Snowflake and Oracle all
  default to SKIP PAST LAST ROW. The two differ in whether matches overlap, so
  a query ported from another engine would otherwise silently return different
  rows. Conversion erases the omitted-vs-explicit distinction, so the rewrite
  has to happen here.
- Deferred constructs are rejected at planning time with actionable messages
  rather than silently mis-executing: ALL ROWS PER MATCH, SUBSET, PERMUTE,
  pattern exclusions, WITHIN, aggregates in DEFINE, NULLS FIRST/LAST, and
  ORDER BY / PARTITION BY on expressions (the last of which otherwise fail
  inside SqlToRelConverter with AssertionError or ClassCastException).

Plan and wire format
- MatchNode at plan.proto tag 19, encoding the pattern as a self-contained
  recursive RowPattern with a symbol table rather than a RexCall tree, with
  field numbers reserved for SUBSET and WITHIN.
- PatternFieldRef in expressions.proto, so RexPatternFieldRef can no longer
  degrade to a plain InputRef and produce wrong-but-type-correct results.
- PinotMatchExchangeNodeInsertRule hash-distributes on the PARTITION BY keys
  and prepends them to the sort collation, so rows arrive clustered and the
  operator can match and flush one partition at a time. A missing PARTITION BY
  is rejected by default, since it collapses the table onto one worker.

Runtime
- PatternToNfaCompiler builds an NFA with prioritized transitions (alternation
  in source order; greedy takes the loop edge first, reluctant the exit edge;
  {n,m} via counter registers rather than state unrolling), so depth-first
  traversal yields the SQL:2016 preferred match first.
- MatchOperator evaluates DEFINE predicates over a classifier tape supporting
  PREV/NEXT/FIRST/LAST/CLASSIFIER/MATCH_NUMBER, emits MEASURES, and advances
  per the AFTER MATCH SKIP mode.
- Guardrails throw rather than truncate, since a truncated pattern result is a
  silently wrong one: maxRowsInMatch, maxStepsPerMatchAttempt, and an
  empty-cycle guard.

v1 covers PARTITION BY, mandatory ORDER BY, MEASURES, ONE ROW PER MATCH, all
four AFTER MATCH SKIP modes, the full pattern algebra including reluctant
quantifiers and anchors, and single-variable aggregates in MEASURES.

MATCH_RECOGNIZE is not yet supported under the v2 physical optimizer or lite
mode; queries there are covered by ignore flags rather than silently wrong
results.
xiangfu0 added a commit to xiangfu0/pinot that referenced this pull request Sep 5, 2026
Adds row pattern recognition to the multi-stage query engine, following the
same shape as the UNNEST support added in apache#17168: a new plan node, an
exchange-insertion rule, and an intermediate-stage operator.

Pinot's parser already accepted the full MATCH_RECOGNIZE grammar (Parser.jj
carries Calcite's Babel production). What was missing was operator-table
registration, validation, a plan node, and all of execution.

Front end
- Register PREV/NEXT/FIRST/LAST/CLASSIFIER/MATCH_NUMBER/RUNNING/FINAL in
  PinotOperatorTable, which is a strict allow-list.
- New MatchRecognizeValidator runs on the SqlNode tree before conversion. It
  rewrites an OMITTED AFTER MATCH clause to SKIP PAST LAST ROW: Calcite
  substitutes SKIP TO NEXT ROW, but SQL:2016, Trino, Snowflake and Oracle all
  default to SKIP PAST LAST ROW. The two differ in whether matches overlap, so
  a query ported from another engine would otherwise silently return different
  rows. Conversion erases the omitted-vs-explicit distinction, so the rewrite
  has to happen here.
- Deferred constructs are rejected at planning time with actionable messages
  rather than silently mis-executing: ALL ROWS PER MATCH, SUBSET, PERMUTE,
  pattern exclusions, WITHIN, aggregates in DEFINE, NULLS FIRST/LAST, and
  ORDER BY / PARTITION BY on expressions (the last of which otherwise fail
  inside SqlToRelConverter with AssertionError or ClassCastException).

Plan and wire format
- MatchNode at plan.proto tag 19, encoding the pattern as a self-contained
  recursive RowPattern with a symbol table rather than a RexCall tree, with
  field numbers reserved for SUBSET and WITHIN.
- PatternFieldRef in expressions.proto, so RexPatternFieldRef can no longer
  degrade to a plain InputRef and produce wrong-but-type-correct results.
- PinotMatchExchangeNodeInsertRule hash-distributes on the PARTITION BY keys
  and prepends them to the sort collation, so rows arrive clustered and the
  operator can match and flush one partition at a time. A missing PARTITION BY
  is rejected by default, since it collapses the table onto one worker.

Runtime
- PatternToNfaCompiler builds an NFA with prioritized transitions (alternation
  in source order; greedy takes the loop edge first, reluctant the exit edge;
  {n,m} via counter registers rather than state unrolling), so depth-first
  traversal yields the SQL:2016 preferred match first.
- MatchOperator evaluates DEFINE predicates over a classifier tape supporting
  PREV/NEXT/FIRST/LAST/CLASSIFIER/MATCH_NUMBER, emits MEASURES, and advances
  per the AFTER MATCH SKIP mode.
- Guardrails throw rather than truncate, since a truncated pattern result is a
  silently wrong one: maxRowsInMatch, maxStepsPerMatchAttempt, and an
  empty-cycle guard.

v1 covers PARTITION BY, mandatory ORDER BY, MEASURES, ONE ROW PER MATCH, all
four AFTER MATCH SKIP modes, the full pattern algebra including reluctant
quantifiers and anchors, and single-variable aggregates in MEASURES.

MATCH_RECOGNIZE is not yet supported under the v2 physical optimizer or lite
mode; queries there are covered by ignore flags rather than silently wrong
results.
xiangfu0 added a commit to xiangfu0/pinot that referenced this pull request Sep 17, 2026
Adds row pattern recognition to the multi-stage query engine, following the
same shape as the UNNEST support added in apache#17168: a new plan node, an
exchange-insertion rule, and an intermediate-stage operator.

Pinot's parser already accepted the full MATCH_RECOGNIZE grammar (Parser.jj
carries Calcite's Babel production). What was missing was operator-table
registration, validation, a plan node, and all of execution.

Front end
- Register PREV/NEXT/FIRST/LAST/CLASSIFIER/MATCH_NUMBER/RUNNING/FINAL in
  PinotOperatorTable, which is a strict allow-list.
- New MatchRecognizeValidator runs on the SqlNode tree before conversion. It
  rewrites an OMITTED AFTER MATCH clause to SKIP PAST LAST ROW: Calcite
  substitutes SKIP TO NEXT ROW, but SQL:2016, Trino, Snowflake and Oracle all
  default to SKIP PAST LAST ROW. The two differ in whether matches overlap, so
  a query ported from another engine would otherwise silently return different
  rows. Conversion erases the omitted-vs-explicit distinction, so the rewrite
  has to happen here.
- Deferred constructs are rejected at planning time with actionable messages
  rather than silently mis-executing: ALL ROWS PER MATCH, SUBSET, PERMUTE,
  pattern exclusions, WITHIN, aggregates in DEFINE, NULLS FIRST/LAST, and
  ORDER BY / PARTITION BY on expressions (the last of which otherwise fail
  inside SqlToRelConverter with AssertionError or ClassCastException).

Plan and wire format
- MatchNode at plan.proto tag 19, encoding the pattern as a self-contained
  recursive RowPattern with a symbol table rather than a RexCall tree, with
  field numbers reserved for SUBSET and WITHIN.
- PatternFieldRef in expressions.proto, so RexPatternFieldRef can no longer
  degrade to a plain InputRef and produce wrong-but-type-correct results.
- PinotMatchExchangeNodeInsertRule hash-distributes on the PARTITION BY keys
  and prepends them to the sort collation, so rows arrive clustered and the
  operator can match and flush one partition at a time. A missing PARTITION BY
  is rejected by default, since it collapses the table onto one worker.

Runtime
- PatternToNfaCompiler builds an NFA with prioritized transitions (alternation
  in source order; greedy takes the loop edge first, reluctant the exit edge;
  {n,m} via counter registers rather than state unrolling), so depth-first
  traversal yields the SQL:2016 preferred match first.
- MatchOperator evaluates DEFINE predicates over a classifier tape supporting
  PREV/NEXT/FIRST/LAST/CLASSIFIER/MATCH_NUMBER, emits MEASURES, and advances
  per the AFTER MATCH SKIP mode.
- Guardrails throw rather than truncate, since a truncated pattern result is a
  silently wrong one: maxRowsInMatch, maxStepsPerMatchAttempt, and an
  empty-cycle guard.

v1 covers PARTITION BY, mandatory ORDER BY, MEASURES, ONE ROW PER MATCH, all
four AFTER MATCH SKIP modes, the full pattern algebra including reluctant
quantifiers and anchors, and single-variable aggregates in MEASURES.

MATCH_RECOGNIZE is not yet supported under the v2 physical optimizer or lite
mode; queries there are covered by ignore flags rather than silently wrong
results.
xiangfu0 added a commit to xiangfu0/pinot that referenced this pull request Sep 19, 2026
Adds row pattern recognition to the multi-stage query engine, following the
same shape as the UNNEST support added in apache#17168: a new plan node, an
exchange-insertion rule, and an intermediate-stage operator.

Pinot's parser already accepted the full MATCH_RECOGNIZE grammar (Parser.jj
carries Calcite's Babel production). What was missing was operator-table
registration, validation, a plan node, and all of execution.

Front end
- Register PREV/NEXT/FIRST/LAST/CLASSIFIER/MATCH_NUMBER/RUNNING/FINAL in
  PinotOperatorTable, which is a strict allow-list.
- New MatchRecognizeValidator runs on the SqlNode tree before conversion. It
  rewrites an OMITTED AFTER MATCH clause to SKIP PAST LAST ROW: Calcite
  substitutes SKIP TO NEXT ROW, but SQL:2016, Trino, Snowflake and Oracle all
  default to SKIP PAST LAST ROW. The two differ in whether matches overlap, so
  a query ported from another engine would otherwise silently return different
  rows. Conversion erases the omitted-vs-explicit distinction, so the rewrite
  has to happen here.
- Deferred constructs are rejected at planning time with actionable messages
  rather than silently mis-executing: ALL ROWS PER MATCH, SUBSET, PERMUTE,
  pattern exclusions, WITHIN, aggregates in DEFINE, NULLS FIRST/LAST, and
  ORDER BY / PARTITION BY on expressions (the last of which otherwise fail
  inside SqlToRelConverter with AssertionError or ClassCastException).

Plan and wire format
- MatchNode at plan.proto tag 19, encoding the pattern as a self-contained
  recursive RowPattern with a symbol table rather than a RexCall tree, with
  field numbers reserved for SUBSET and WITHIN.
- PatternFieldRef in expressions.proto, so RexPatternFieldRef can no longer
  degrade to a plain InputRef and produce wrong-but-type-correct results.
- PinotMatchExchangeNodeInsertRule hash-distributes on the PARTITION BY keys
  and prepends them to the sort collation, so rows arrive clustered and the
  operator can match and flush one partition at a time. A missing PARTITION BY
  is rejected by default, since it collapses the table onto one worker.

Runtime
- PatternToNfaCompiler builds an NFA with prioritized transitions (alternation
  in source order; greedy takes the loop edge first, reluctant the exit edge;
  {n,m} via counter registers rather than state unrolling), so depth-first
  traversal yields the SQL:2016 preferred match first.
- MatchOperator evaluates DEFINE predicates over a classifier tape supporting
  PREV/NEXT/FIRST/LAST/CLASSIFIER/MATCH_NUMBER, emits MEASURES, and advances
  per the AFTER MATCH SKIP mode.
- Guardrails throw rather than truncate, since a truncated pattern result is a
  silently wrong one: maxRowsInMatch, maxStepsPerMatchAttempt, and an
  empty-cycle guard.

v1 covers PARTITION BY, mandatory ORDER BY, MEASURES, ONE ROW PER MATCH, all
four AFTER MATCH SKIP modes, the full pattern algebra including reluctant
quantifiers and anchors, and single-variable aggregates in MEASURES.

MATCH_RECOGNIZE is not yet supported under the v2 physical optimizer or lite
mode; queries there are covered by ignore flags rather than silently wrong
results.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
xiangfu0 added a commit to xiangfu0/pinot that referenced this pull request Sep 30, 2026
Adds row pattern recognition to the multi-stage query engine, following the
same shape as the UNNEST support added in apache#17168: a new plan node, an
exchange-insertion rule, and an intermediate-stage operator.

Pinot's parser already accepted the full MATCH_RECOGNIZE grammar (Parser.jj
carries Calcite's Babel production). What was missing was operator-table
registration, validation, a plan node, and all of execution.

Front end
- Register PREV/NEXT/FIRST/LAST/CLASSIFIER/MATCH_NUMBER/RUNNING/FINAL in
  PinotOperatorTable, which is a strict allow-list.
- New MatchRecognizeValidator runs on the SqlNode tree before conversion. It
  rewrites an OMITTED AFTER MATCH clause to SKIP PAST LAST ROW: Calcite
  substitutes SKIP TO NEXT ROW, but SQL:2016, Trino, Snowflake and Oracle all
  default to SKIP PAST LAST ROW. The two differ in whether matches overlap, so
  a query ported from another engine would otherwise silently return different
  rows. Conversion erases the omitted-vs-explicit distinction, so the rewrite
  has to happen here.
- Deferred constructs are rejected at planning time with actionable messages
  rather than silently mis-executing: ALL ROWS PER MATCH, SUBSET, PERMUTE,
  pattern exclusions, WITHIN, aggregates in DEFINE, NULLS FIRST/LAST, and
  ORDER BY / PARTITION BY on expressions (the last of which otherwise fail
  inside SqlToRelConverter with AssertionError or ClassCastException).

Plan and wire format
- MatchNode at plan.proto tag 19, encoding the pattern as a self-contained
  recursive RowPattern with a symbol table rather than a RexCall tree, with
  field numbers reserved for SUBSET and WITHIN.
- PatternFieldRef in expressions.proto, so RexPatternFieldRef can no longer
  degrade to a plain InputRef and produce wrong-but-type-correct results.
- PinotMatchExchangeNodeInsertRule hash-distributes on the PARTITION BY keys
  and prepends them to the sort collation, so rows arrive clustered and the
  operator can match and flush one partition at a time. A missing PARTITION BY
  is rejected by default, since it collapses the table onto one worker.

Runtime
- PatternToNfaCompiler builds an NFA with prioritized transitions (alternation
  in source order; greedy takes the loop edge first, reluctant the exit edge;
  {n,m} via counter registers rather than state unrolling), so depth-first
  traversal yields the SQL:2016 preferred match first.
- MatchOperator evaluates DEFINE predicates over a classifier tape supporting
  PREV/NEXT/FIRST/LAST/CLASSIFIER/MATCH_NUMBER, emits MEASURES, and advances
  per the AFTER MATCH SKIP mode.
- Guardrails throw rather than truncate, since a truncated pattern result is a
  silently wrong one: maxRowsInMatch, maxStepsPerMatchAttempt, and an
  empty-cycle guard.

v1 covers PARTITION BY, mandatory ORDER BY, MEASURES, ONE ROW PER MATCH, all
four AFTER MATCH SKIP modes, the full pattern algebra including reluctant
quantifiers and anchors, and single-variable aggregates in MEASURES.

MATCH_RECOGNIZE is not yet supported under the v2 physical optimizer or lite
mode; queries there are covered by ignore flags rather than silently wrong
results.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

join Related to JOIN operations multi-stage Related to the multi-stage query engine

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants