Skip to content

Multi-Stage: Add SQL:2016 MATCH_RECOGNIZE (row pattern recognition) - #19311

Open
xiangfu0 wants to merge 7 commits into
apache:masterfrom
xiangfu0:claude/match-recognize-oss
Open

xiangfu0 wants to merge 7 commits into
apache:masterfrom
xiangfu0:claude/match-recognize-oss

Conversation

@xiangfu0

@xiangfu0 xiangfu0 commented Aug 19, 2026 •

Copy link
Copy Markdown
Contributor

PR flow

Adds MATCH_RECOGNIZE validation, planning, exchange insertion, serialization, deserialization, and runtime execution with limits and tape-based matching.

flowchart TD
  N0["Validation #40;F30#41;"]:::stAdded
  N1["PlanNode Construction #40;F15#41;"]:::stAdded
  N2["Exchange Insertion #40;F4#41;"]:::stAdded
  N3["Serialization #40;F26#41;"]:::stAdded
  N4["Deserialization #40;F25#41;"]:::stAdded
  N5["Runtime Support #40;F33#44; F35#44; F37#41;"]:::stAdded
  N0 -->|"feeds validated SQL to planner"| N1
  N1 -->|"MatchNode triggers exchange rule"| N2
  N2 -->|"plan with exchange serialized"| N3
  N3 -->|"protobuf sent over wire"| N4
  N4 -->|"MatchNode drives matcher"| N5
  classDef stAdded fill:#dafbe1,stroke:#1a7f37,color:#1f2328,stroke-width:2px
  classDef stModified fill:#fff8c5,stroke:#9a6700,color:#1f2328,stroke-width:2px
  classDef stRemoved fill:#ffebe9,stroke:#cf222e,color:#1f2328,stroke-width:2px
  classDef stUnchanged fill:#f6f8fa,stroke:#656d76,color:#1f2328,stroke-width:1px
Loading

AI-generated · Green: added · Yellow: modified · Red: removed · Gray: existing

Partial evidence: 29 file patches omitted; 2 truncated.

Diff evidence
  • F4: pinot-query-planner/src/main/java/org/apache/pinot/calcite/rel/rules/PinotMatchExchangeNodeInsertRule.java — after
  • F15: pinot-query-planner/src/main/java/org/apache/pinot/query/planner/logical/RelToPlanNodeConverter.java — before · after
  • F25: pinot-query-planner/src/main/java/org/apache/pinot/query/planner/serde/PlanNodeDeserializer.java — before · after
  • F26: pinot-query-planner/src/main/java/org/apache/pinot/query/planner/serde/PlanNodeSerializer.java — before · after
  • F30: pinot-query-planner/src/main/java/org/apache/pinot/query/validate/MatchRecognizeValidator.java — after
  • F33: pinot-query-runtime/src/main/java/org/apache/pinot/query/runtime/QueryRunner.java — before · after
  • F35: pinot-query-runtime/src/main/java/org/apache/pinot/query/runtime/operator/MultiStageOperator.java — before · after
  • F37: pinot-query-runtime/src/main/java/org/apache/pinot/query/runtime/operator/match/MatchLimits.java — after
  • Regenerate PR flow

Adds SQL:2016 row pattern recognition to Pinot's multi-stage query engine: validation, logical/wire plans, exchange planning, and an intermediate-stage matcher.

Pinot's parser already accepts the MATCH_RECOGNIZE grammar through Calcite's Babel parser. This PR adds the missing supported-subset validation and execution path.

Design and delivery boundary

The parser validation, MatchNode wire representation, exchange rule, runtime operator, and expression semantics form one end-to-end query operator and must land atomically. Splitting those layers would either expose syntax that cannot execute or dispatch plans that servers cannot interpret. Existing non-MATCH_RECOGNIZE queries keep their existing plan and execution behavior; unsupported MATCH_RECOGNIZE modes fail closed with actionable validation errors. The PR remains labeled design-review for maintainer sign-off before landing.

Rollback is likewise atomic: revert the feature commits and stop issuing MATCH_RECOGNIZE queries. There is no persisted-data or table-schema migration.

Front end and planner

  • Registers PREV/NEXT/FIRST/LAST/CLASSIFIER/MATCH_NUMBER/RUNNING/FINAL in the strict PinotOperatorTable allow-list.
  • Adds MatchRecognizeValidator, including SQL:2016's omitted AFTER MATCH default (SKIP PAST LAST ROW) and actionable rejection of deferred constructs such as ALL ROWS PER MATCH, SUBSET, PERMUTE, exclusions, WITHIN, aggregates in DEFINE, explicit null ordering, expression partition/order keys, and multi-value partition or aggregate inputs.
  • Applies SQL identifier rules consistently: unquoted pattern variables are case-insensitive, quoted identifiers retain case, and row-source aliases remain isolated from the pattern-variable namespace.
  • Treats RUNNING in DEFINE as the clause's intrinsic mode and rejects unsupported FINAL semantics there.
  • Derives CLASSIFIER() as nullable so Calcite preserves COUNT(classifier_measure) instead of incorrectly rewriting it to COUNT(*) for empty matches.
  • Adds MatchNode and a self-contained recursive RowPattern; PatternFieldRef preserves pattern-variable identity instead of degrading to a plain input reference.
  • Hash-distributes by PARTITION BY, then sorts by the exchange's actual partition-key order plus the requested order keys. Queries without PARTITION BY require the explicit allowMatchRecognizeWithoutPartitionBy query option and execute as one global partition.
  • The v2 physical optimizer and lite mode are not supported yet. The planner now checks the effective broker defaults plus query overrides and rejects these combinations explicitly; SET usePhysicalOptimizer = false selects the supported path.
  • Sender-side sorting for sort exchanges is tracked separately in [MSE] Support sender-side sorting for sort exchanges #19395.

Runtime and resource safety

  • Compiles prioritized NFAs for source-order alternation, greedy/reluctant quantifiers, bounded quantifiers, anchors, and quantified zero-width anchors without empty-cycle hangs.
  • Evaluates DEFINE predicates and MEASURES over a classifier tape, including navigation around CLASSIFIER(), bounded navigation, and exact decimal/integer aggregation behavior.
  • Emits empty matches once at each input position and always advances one row for all supported skip modes. Missing skip targets remain errors for non-empty matches.
  • Traverses the contiguous universal match range directly for aggregates, avoiding per-match row-index materialization; pattern-variable aggregates retain their sparse symbol indexes.
  • Processes one monotonically ordered partition at a time without retaining an unbounded closed-partition set or allocating a key tuple per row.
  • Emits at most 1,024 rows per transfer block and resumes matching state across calls; upstream error blocks supersede buffered output.
  • Checks cancellation/deadline state every 1,024 matcher transitions.
  • Stores backtracking state in primitive choice/counter logs and enforces a non-configurable 64 MiB cap on their combined array payload before growth.
  • Throws instead of truncating when a row, step, or retained-state limit is reached.

Resource-limit precedence is node hint, query option, server config, then default:

Limit Query option Hint key Server config
Rows buffered per partition maxRowsInMatchPartition max_rows_in_match_partition pinot.query.match.max.rows.per.partition
NFA transitions per start position maxStepsPerMatchAttempt max_steps_per_match_attempt pinot.query.match.max.steps.per.attempt

Scope

v1 covers PARTITION BY, mandatory ORDER BY, MEASURES, ONE ROW PER MATCH, all four supported AFTER MATCH SKIP modes, pattern alternation/concatenation/quantifiers/anchors, navigation and classifier functions, and single-variable aggregates in measures.

Trino comparison

The supported v1 behavior and tests were compared against Trino at commit 19b4cebac145, especially its row-pattern query suite, aggregation suite, analyzer suite, and feature documentation.

Coverage added from that comparison includes:

  • empty-match advancement for every supported skip mode, empty-match CLASSIFIER() nullability, and COUNT(classifier_measure) versus COUNT(*) semantics;
  • FIRST/LAST/PREV/NEXT navigation around CLASSIFIER(), including out-of-match nulls and composition of non-zero logical and physical offsets;
  • intrinsically running DEFINE expressions and rejection of FINAL in DEFINE;
  • case-insensitive unquoted pattern variables, case-sensitive quoted variables, and separation of row-source aliases from the pattern-variable namespace;
  • allocation-free universal aggregates over bounded mixed-null ranges, per-variable aggregates over empty/all-null/mixed-null input and repeated sparse pattern-variable rows, plus empty input; and
  • deterministic multi-key ordering with mixed DESC/ASC keys.

This comparison targets Pinot's explicitly supported v1 subset and does not claim full Trino feature parity. ALL ROWS PER MATCH, SUBSET, PERMUTE, exclusions, WITHIN, aggregates in DEFINE, optional measures/order, and row-pattern window frames remain deferred and fail validation where applicable.

Rolling-upgrade boundary

MatchNode is plan oneof field 19. Older servers preserve that protobuf field as unknown but see the node oneof as unset (NODE_NOT_SET), so they cannot execute a dispatched MATCH_RECOGNIZE stage. Upgrade all servers before issuing these queries; there is no mixed-version fallback. Existing query node kinds remain wire-compatible and are unaffected.

Testing

  • At 35692e0aad54, 726 targeted tests across 20 common, planner, runtime, segment, upsert, and compression classes passed with zero failures, errors, or skips. This includes the restored query-option validation, MATCH expression/limit tests, plan merge/equivalence tests, and the current upstream string-function behavior.
  • MatchRecognizeIntegrationTest: 21/21 passed against a real two-server cluster using the coherently compiled and installed artifacts from this source.
  • The dependency build and main/test compilation passed for all 65 reactor projects selected by the affected modules. Spotless, Checkstyle, license formatting/checking, and git diff --check passed for the affected modules and root on JDK 25.
  • Independent review verified all 65 MATCH feature files against the recovered reviewed commits and current upstream, plus the current factory APIs and dependency versions in the built artifacts. There are no compiler warnings on MATCH-owned added lines; three imported upstream test-constructor warnings remain unchanged from master.
  • Fresh hosted CI passed on 35692e0aad54: both unit suites, both integration suites, linter, Quickstart, binary compatibility, all four compatibility regressions, Java 11 client compatibility, Docker and Trivy. Earlier results from 80f23878ef6d remain historical.

@codecov-commenter

codecov-commenter commented Aug 19, 2026 •

Copy link
Copy Markdown

Codecov Report

✅ All modified and coverable lines are covered by tests.
✅ Project coverage is 0.00%. Comparing base (7a00b23) to head (35692e0).
⚠️ Report is 2 commits behind head on master.

❗ There is a different number of reports uploaded between BASE (7a00b23) and HEAD (35692e0). Click for more details.

HEAD has 14 uploads less than BASE
Flag BASE (7a00b23) HEAD (35692e0)
unittests2 1 0
unittests 1 0
java-25 5 2
temurin 5 2
lane-a 2 1
integration 4 2
integration1 2 0
lane-b 2 1
Additional details and impacted files
@@              Coverage Diff              @@
##             master   #19311       +/-   ##
=============================================
- Coverage     39.94%    0.00%   -39.95%     
=============================================
  Files          3520        3     -3517     
  Lines        228940        6   -228934     
  Branches      36313        0    -36313     
=============================================
- Hits          91453        0    -91453     
+ Misses       129215        6   -129209     
+ Partials       8272        0     -8272     
Flag Coverage Δ
integration 0.00% <ø> (-100.00%) ⬇️
integration1 ?
integration2 0.00% <ø> (ø)
java-25 0.00% <ø> (-39.95%) ⬇️
lane-a 0.00% <ø> (-100.00%) ⬇️
lane-b 0.00% <ø> (ø)
temurin 0.00% <ø> (-39.95%) ⬇️
unittests ?
unittests2 ?

Flags with carried forward coverage won't be shown. Click here to find out more.

☔ View full report in Codecov by Harness.
📢 Have feedback on the report? Share it here.

🚀 New features to boost your workflow:
  • 📦 JS Bundle Analysis: Save yourself from yourself by tracking and limiting bundle sizes in JS merges.

@xiangfu0 xiangfu0 added feature New functionality multi-stage Related to the multi-stage query engine query Related to query processing sql-compliance Related to SQL standard compliance configuration Config changes (addition/deletion/change in behavior) release-notes Referenced by PRs that need attention when compiling the next release notes labels Aug 19, 2026
@xiangfu0
xiangfu0 force-pushed the claude/match-recognize-oss branch from ededa2b to 3435400 Compare August 20, 2026 09:03
Comment thread pinot-query-runtime/src/test/resources/queries/MatchRecognize.json
@xiangfu0
xiangfu0 force-pushed the claude/match-recognize-oss branch 3 times, most recently from 7ebcfdc to 5f27d70 Compare August 26, 2026 09:21

@xiangfu0 xiangfu0 left a comment

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Review of current head 5f27d70c.

The correctness, rolling-upgrade, and resource-behavior issues below should be addressed before merge. I did not duplicate the two existing unresolved threads about v2/lite rejection and _closedPartitionKeys cardinality.

Process follow-ups:

  • This is a 56-file, 7.5K-line change spanning protobuf, planning, execution, configuration, and integration. Please either link the reviewed design/maintainer agreement for keeping it as one vertical slice, or split it into reviewable stacked PRs.
  • Please link the sender-side-sorting TODO to a tracking issue.
  • Both commits contain AI Co-Authored-By trailers; repository guidance says to omit those, so please rewrite the commit messages before merge.

Comment thread pinot-common/src/main/proto/plan.proto
@xiangfu0
xiangfu0 force-pushed the claude/match-recognize-oss branch 4 times, most recently from 9623e5c to ccb8dea Compare August 29, 2026 09:07
@xiangfu0 xiangfu0 added design-review Requires design review before implementation backward-incompat Introduces a backward-incompatible API or behavior change upgrade-incompat PR may introduce incompatibility during upgrade of an installation labels Aug 29, 2026
@xiangfu0
xiangfu0 force-pushed the claude/match-recognize-oss branch from ccb8dea to b612fab Compare August 29, 2026 10:05
@xiangfu0

Copy link
Copy Markdown
Contributor Author

Addressed the process follow-ups on head b612fab43e:

  • Expanded the PR description with the end-to-end design/atomicity rationale, rollback boundary, exact supported/unsupported modes, resource-limit precedence, and the required all-servers-first rolling-upgrade sequence. The design-review label remains so maintainer sign-off is still explicit before landing.
  • Linked the sender-side sort TODO to [MSE] Support sender-side sorting for sort exchanges #19395.
  • Rewrote the feature commits to remove both AI Co-Authored-By trailers.
  • Replied to and resolved all 13 inline threads with the pushed fix and regression evidence.

Local validation passed: the 342-test focused suite, a 32-test matcher/operator rerun after the final retained-state changes, 13/13 two-server integration cases, the full 63-module test-compile reactor, Spotless, checkstyle, license checks, and diff hygiene. Fresh CI for this rewritten head is running.

@xiangfu0
xiangfu0 force-pushed the claude/match-recognize-oss branch 6 times, most recently from e25bf6f to c05a54b Compare September 5, 2026 00:12
@xiangfu0

xiangfu0 commented Sep 5, 2026 •

Copy link
Copy Markdown
Contributor Author

The rebased PR head is now 3034c6d21fef on upstream master at e6787e178645. The production tree is unchanged from the previously green c05a54b7d69b; this force-push adds a test-only follow-up from the pinned Trino comparison.

Coverage added in this follow-up:

  • Field-by-field positive/negative tests for PlanNodeMerger.visitMatch and EquivalentStagesFinder.visitMatch, covering pattern definitions, pattern shape, measures, partitioning, ordering, skip mode/target, rows-per-match, node type, and input/base mismatch.
  • The positive merger case observes recursively merged explain attributes (rows=2+3=5), so it verifies that the merged child is actually installed.
  • Repeated-pattern-variable coverage for PATTERN (A B A) verifies that aggregates and logical navigation use sparse classifier rows rather than the contiguous match span.
  • Runtime expression cases for default FIRST/LAST/PREV/NEXT offsets, nested PREV(RUNNING LAST(...)), composition of non-zero logical and physical navigation offsets, out-of-range nulls, SQL true/false/unknown DEFINE truth semantics, all-null aggregate identities, and defensive stored-type conversion branches.

The comparison remains pinned to Trino 19b4cebac145. It targets Pinot's supported v1 subset and does not claim full Trino feature or planner type parity.

Validation:

  • Expanded focused common/planner/runtime matrix: 434 invocations; 422 passed and 12 expected lite/v2-optimizer variants skipped
  • Planner merge/equivalence suites: 47/47; final strengthened PlanNodeMergerTest rerun: 14/14
  • Core runtime matcher/expression/NFA suites: 74/74
  • Two-server MatchRecognizeIntegrationTest: 21/21 on the unchanged production tree through the full 63-module JDK 25 reactor
  • Local JaCoCo: all 23 executable lines in PlanNodeMerger.visitMatch and all 12 in EquivalentStagesFinder.visitMatch covered, with no missed or partial lines
  • Spotless, checkstyle, license format/check, and git diff --check: clean after the final test change
  • Independent final review: no remaining blockers or Trino-v1 test gaps for this follow-up

Fresh hosted CI is complete: 12/12 checks passed for 3034c6d21fef. Final Codecov reports 81.91882% patch coverage with 343 changed lines missing and 67.86% project coverage (+0.16% versus base); all expected uploads are present and no coverage flag is unknown.

@xiangfu0
xiangfu0 force-pushed the claude/match-recognize-oss branch 2 times, most recently from 6575e4c to 3034c6d Compare September 5, 2026 01:54
@xiangfu0
xiangfu0 force-pushed the claude/match-recognize-oss branch 2 times, most recently from b6ed53c to d252451 Compare September 19, 2026 01:33
xiangfu0 and others added 2 commits September 30, 2026 10:33
Adds row pattern recognition to the multi-stage query engine, following the
same shape as the UNNEST support added in apache#17168: a new plan node, an
exchange-insertion rule, and an intermediate-stage operator.

Pinot's parser already accepted the full MATCH_RECOGNIZE grammar (Parser.jj
carries Calcite's Babel production). What was missing was operator-table
registration, validation, a plan node, and all of execution.

Front end
- Register PREV/NEXT/FIRST/LAST/CLASSIFIER/MATCH_NUMBER/RUNNING/FINAL in
  PinotOperatorTable, which is a strict allow-list.
- New MatchRecognizeValidator runs on the SqlNode tree before conversion. It
  rewrites an OMITTED AFTER MATCH clause to SKIP PAST LAST ROW: Calcite
  substitutes SKIP TO NEXT ROW, but SQL:2016, Trino, Snowflake and Oracle all
  default to SKIP PAST LAST ROW. The two differ in whether matches overlap, so
  a query ported from another engine would otherwise silently return different
  rows. Conversion erases the omitted-vs-explicit distinction, so the rewrite
  has to happen here.
- Deferred constructs are rejected at planning time with actionable messages
  rather than silently mis-executing: ALL ROWS PER MATCH, SUBSET, PERMUTE,
  pattern exclusions, WITHIN, aggregates in DEFINE, NULLS FIRST/LAST, and
  ORDER BY / PARTITION BY on expressions (the last of which otherwise fail
  inside SqlToRelConverter with AssertionError or ClassCastException).

Plan and wire format
- MatchNode at plan.proto tag 19, encoding the pattern as a self-contained
  recursive RowPattern with a symbol table rather than a RexCall tree, with
  field numbers reserved for SUBSET and WITHIN.
- PatternFieldRef in expressions.proto, so RexPatternFieldRef can no longer
  degrade to a plain InputRef and produce wrong-but-type-correct results.
- PinotMatchExchangeNodeInsertRule hash-distributes on the PARTITION BY keys
  and prepends them to the sort collation, so rows arrive clustered and the
  operator can match and flush one partition at a time. A missing PARTITION BY
  is rejected by default, since it collapses the table onto one worker.

Runtime
- PatternToNfaCompiler builds an NFA with prioritized transitions (alternation
  in source order; greedy takes the loop edge first, reluctant the exit edge;
  {n,m} via counter registers rather than state unrolling), so depth-first
  traversal yields the SQL:2016 preferred match first.
- MatchOperator evaluates DEFINE predicates over a classifier tape supporting
  PREV/NEXT/FIRST/LAST/CLASSIFIER/MATCH_NUMBER, emits MEASURES, and advances
  per the AFTER MATCH SKIP mode.
- Guardrails throw rather than truncate, since a truncated pattern result is a
  silently wrong one: maxRowsInMatch, maxStepsPerMatchAttempt, and an
  empty-cycle guard.

v1 covers PARTITION BY, mandatory ORDER BY, MEASURES, ONE ROW PER MATCH, all
four AFTER MATCH SKIP modes, the full pattern algebra including reluctant
quantifiers and anchors, and single-variable aggregates in MEASURES.

MATCH_RECOGNIZE is not yet supported under the v2 physical optimizer or lite
mode; queries there are covered by ignore flags rather than silently wrong
results.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Upstream now enforces `///` markdown doc comments (JEP 467) over `/** */`
Javadoc via a checkstyle RegexpCheck, and this feature branch predates that
rule. Converts all 168 Javadoc blocks across the 23 MATCH_RECOGNIZE files, and
removes an `org.apache.calcite.sql.SqlLiteral` import that the feature commit
added without ever using (UnusedImports flags it independently).

Formatting only: `{@link X}` becomes `[X]`, `{@link X label}` becomes
`[label][X]`, `{@code x}` becomes a backtick span, `<p>` becomes an empty ///
line, and the remaining HTML becomes its markdown equivalent. No documentation
text was reworded or dropped, and the removed import is the only non-comment
line changed.

Verified: checkstyle reports 0 violations in each of pinot-spi, pinot-common,
pinot-query-planner, pinot-query-runtime and pinot-integration-tests (run
per-module, since a combined reactor stops at the first failure and hides the
rest); `javadoc -Xdoclint:reference,syntax` is clean over the changed sources;
pinot-query-planner 1598 tests and pinot-query-runtime 4611 tests still pass.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
@xiangfu0
xiangfu0 force-pushed the claude/match-recognize-oss branch from d252451 to 80f2387 Compare September 30, 2026 17:41
xiangfu0 and others added 5 commits October 4, 2026 19:44
Harden planner validation, configuration precedence, wire compatibility, partition ordering, matcher resource bounds, cancellation, and output buffering. Add focused and two-server integration coverage for null semantics, global execution, legacy plans, numeric precision, and retained matcher state.

(cherry picked from commit 6aa6d26)
@xiangfu0

xiangfu0 commented Oct 5, 2026 •

Copy link
Copy Markdown
Contributor Author

Recovered the two reviewed MATCH hardening/configuration commits that had been dropped from this branch, and integrated current master f5d9be624486, in 35692e0aad54. The restored changes include canonical typed query options, validation and limit precedence, resource/cancellation guards, bounded output and partition release, RUNNING/FINAL behavior, identifier isolation, and the missing regression tests. The field-19 rolling-upgrade boundary and explicit optimizer exclusions remain documented.

Current upstream also needs a direct managed, provided JetBrains annotations dependency for segment-local compilation: zstd-jni's provided annotation dependency is not transitive. This changes compile-time dependency declaration only.

Validation on this source: 726 targeted tests and all 21 real two-server MATCH integration cases passed with zero failures, errors, or skips. The affected-module/root style and license checks and independent source/artifact review passed. Fresh hosted CI also passed on this exact head: both unit suites, both integration suites, linter, Quickstart, binary compatibility, all four compatibility regressions, Java 11 client compatibility, Docker and Trivy. The PR description records these current results separately from historical results.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

backward-incompat Introduces a backward-incompatible API or behavior change configuration Config changes (addition/deletion/change in behavior) design-review Requires design review before implementation feature New functionality multi-stage Related to the multi-stage query engine query Related to query processing release-notes Referenced by PRs that need attention when compiling the next release notes sql-compliance Related to SQL standard compliance upgrade-incompat PR may introduce incompatibility during upgrade of an installation

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants