Repository navigation
perf(sql): reduce SQL planning allocations for large literal arrays - #20367
FrankChen021 wants to merge 5 commits into
Conversation
There was a problem hiding this comment.
🟢 Approval recommended
The optimization is type-gated, preserves fallback evaluation, and includes focused correctness and benchmark coverage.
Pull request overview
Reduces SQL planning allocations for large literal string and integer arrays, especially arrays generated from large IN lists.
Changes:
- Adds a type-safe literal-array reuse fast path.
- Adds planner and SQL semantic regression tests.
- Adds string and long plan-only benchmark variants.
File summaries
| File | Description |
|---|---|
DruidRexExecutor.java |
Reuses compatible literal array calls. |
DruidRexExecutorTest.java |
Tests type matching and evaluator fallbacks. |
CalciteQueryTest.java |
Verifies IN/NOT IN semantics. |
InPlanningBenchmark.java |
Measures planning-only performance. |
Review details
- Files reviewed: 4/4 changed files
- Comments generated: 0
- Review effort level: Lite
💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.
FrankChen021
left a comment
There was a problem hiding this comment.
🟢 Approval recommended
Static review of the current head found no actionable PR-caused issues. The literal-array fast path is restricted to ARRAY_VALUE_CONSTRUCTOR calls whose component type is CHAR/VARCHAR or TINYINT/SMALLINT/INTEGER/BIGINT and whose operands are RexLiteral values with matching types ignoring nullability; casts, computed operands, mixed or mismatched types, unsupported types, and required normalization continue through the existing Druid evaluator path.
Reviewed 4 of 4 changed files:
- benchmarks/src/test/java/org/apache/druid/benchmark/query/InPlanningBenchmark.java
- sql/src/main/java/org/apache/druid/sql/calcite/planner/DruidRexExecutor.java
- sql/src/test/java/org/apache/druid/sql/calcite/CalciteQueryTest.java
- sql/src/test/java/org/apache/druid/sql/calcite/planner/DruidRexExecutorTest.java
I also inspected the executor registration and callers, scalar-IN and array conversions, Calcite array type/coercion behavior, relevant test data, and the benchmark lifecycle. Validation run: git diff --check against the PR merge base; it passed. No builds or broad tests were run.
This is an automated review by Codex GPT-5.6-Luna(max)
| /// | ||
| /// Returns null for mismatched element types, unsupported types, or non-literal operands (including CAST calls). | ||
| /// These retain the existing Druid expression evaluation path, including any required type normalization. | ||
| /// |
There was a problem hiding this comment.
/// Seems like a wrong JavaDoc format
There was a problem hiding this comment.
This is markdown format Java doc from JDK 23
|
I will keep this open until druid 39 is going to be cut off |
Related to #20326.
Description
Large literal arrays already contain constant Calcite values, but the existing planning path converts them into Druid expressions, formats them as text, parses that text, evaluates it, and rebuilds Calcite literals:
This round trip allocates per-element expression objects, strings, parser objects, evaluated values, and replacement Calcite nodes.
This change skips that work for already-normalized string and integer literal arrays. Expressions requiring conversion or evaluation retain the existing path.
Benchmark results
MB is decimal; allocation means total bytes allocated per planning operation, not peak or retained heap.
Both variants use Calcite 1.42.0 with only the CALCITE-7782 fix. The benchmark isolates this Druid change using identical dependencies and harness settings.
JDK 25.0.4.1; two forks; two warmup and three measurement iterations; GC profiler. Queries are planned without execution or EXPLAIN serialization. Timing results are preliminary.
Tests
Tests cover identity reuse, NULLs, integer boundaries, type mismatches, charset/collation differences, computed operands, and SQL IN/NOT IN semantics.
Release note
Reduced SQL planning allocations for large literal string and integer arrays, including arrays generated from large IN lists.