Skip to content

perf: optimize some multi-value string operators to allow planning to specialized ListFilteredVirtualColumn without extractionFn - #20469

Open
clintropolis wants to merge 1 commit into
apache:masterfrom
clintropolis:optimize-list-filtered-expressions
Open

clintropolis wants to merge 1 commit into
apache:masterfrom
clintropolis:optimize-list-filtered-expressions

Conversation

@clintropolis

Copy link
Copy Markdown
Member

Description

Optimizes SQL planning so that multi-value string operators that would plan to a specialized ListFilteredVirtualColumn were sqlUseExtractionFn set to true, but currently plan to a pure expression, to instead still plan to the specialized column pointed at an expression virtual column. If the expression virtual column is actually a transform on top of a dictionary encoded column, the performance can be approximately the same as the extractionFn version since those selectors can do deferred evaluation and provide access to the underlying dictionary ids.

From the added benchmark:

basic dataset: 1.5M rows, dimMultivalEnumerated has 5 distinct values

# query expression (ms/op) specialized (ms/op) extractionFn (ms/op) speedup
0 GROUP BY MV_FILTER_ONLY(LOOKUP(dimMultivalEnumerated, ...), ARRAY['greeting']) 1072 ± 185 161 ± 52 154 ± 18 6.6x
1 GROUP BY MV_FILTER_NONE(LOOKUP(dimMultivalEnumerated, ...), ARRAY['greeting']) 1098 ± 145 168 ± 25 176 ± 16 6.5x
2 GROUP BY MV_FILTER_REGEX(LOOKUP(dimMultivalEnumerated, ...), '^g.*') 1179 ± 255 156 ± 8 154 ± 18 7.6x
3 GROUP BY MV_FILTER_PREFIX(LOOKUP(dimMultivalEnumerated, ...), 'g') 1141 ± 490 156 ± 20 154 ± 19 7.3x
4 WHERE MV_FILTER_ONLY(LOOKUP(dimMultivalEnumerated, ...), ARRAY['greeting']) = 'greeting' 1060 ± 185 85 ± 7 82 ± 12 12.5x
9 GROUP BY MV_FILTER_ONLY(SUBSTRING(dimMultivalEnumerated, 1, 1), ARRAY['B']) 1069 ± 158 152 ± 42 147 ± 16 7.1x
10 GROUP BY MV_FILTER_ONLY(REGEXP_EXTRACT(dimMultivalEnumerated, '^[A-Z][a-z]'), ARRAY['Ba']) 1086 ± 186 152 ± 29 157 ± 16 7.2x

grouper dataset: 2 segments × 1.5M rows, multi-value string columns with 1M distinct values; the lookup maps 1M keys to 1000 values

# query expression (ms/op) specialized (ms/op) extractionFn (ms/op) speedup
5 GROUP BY MV_FILTER_ONLY(LOOKUP("multi-string-Uniform-1_000_000", ...), ARRAY['bucket-1', 'bucket-2']) 7119 ± 2195 1025 ± 335 981 ± 517 6.9x
6 GROUP BY MV_FILTER_ONLY(LOOKUP("multi-string-ZipF-1_000_000", ...), ARRAY['bucket-1', 'bucket-2']) 5859 ± 606 1192 ± 125 1180 ± 52 4.9x
7 GROUP BY MV_FILTER_PREFIX(LOOKUP("multi-string-Uniform-1_000_000", ...), 'bucket-1') 7893 ± 4515 1351 ± 21 1316 ± 384 5.8x
8 WHERE MV_FILTER_ONLY(LOOKUP("multi-string-Uniform-1_000_000", ...), ARRAY['bucket-1']) = 'bucket-1' 7557 ± 367 887 ± 359 911 ± 214 8.5x
11 GROUP BY MV_FILTER_ONLY(SUBSTRING("multi-string-Uniform-1_000_000", 1, 2), ARRAY['12', '34']) 3858 ± 1815 568 ± 103 597 ± 214 6.8x
12 GROUP BY MV_FILTER_ONLY(REGEXP_EXTRACT("multi-string-Uniform-1_000_000", '[0-9]{2}$'), ARRAY['12', '34']) 5059 ± 5050 796 ± 83 747 ± 22 6.4x

… specialized ListFilteredVirtualColumn without extractionFn
@Override
public String getQuery()
{
return QUERIES.get(Integer.parseInt(query)).rhs;
@Override
public List<String> getDatasources()
{
return Collections.singletonList(QUERIES.get(Integer.parseInt(query)).lhs);

@FrankChen021 FrankChen021 left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🟡 Changes recommended

Reviewed all 9 changed files in the complete merge-base diff, including the benchmark additions, selector and mapping changes, SQL expression conversion, and associated tests.

The PR should not merge until the removed public selector-helper overloads are retained for source and binary compatibility and nondeterministic one-input delegates are excluded from deferred dictionary evaluation. The optimized deterministic paths and duplicate-value mapping were otherwise traced through the virtual-column registry and selector implementations.

Validation: static review only; git diff --check f45885a53b1a16d6930caa30b4077d05e25a5a86 b4adf42c27d2667c1453add04e7e39abf219288b passed. No builds or tests were run per review instructions.

Severity Findings
P0 0
P1 1
P2 1
P3 0
Total 2

This is an automated review by Codex GPT-5.6-Luna(max)

After addressing the findings or replying to the comments, you can request another review from me to trigger a new automated review.

return null;
}
final Expr.BindingAnalysis bindingAnalysis = plannerContext.parseExpression(arg.getExpression()).analyzeInputs();
if (bindingAnalysis.getRequiredBindings().size() != 1

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

[P2] Do not defer nondeterministic delegates by dictionary

Finding: The new eligibility check accepts any one-binding scalar string expression without checking BindingAnalysis.isNonDeterministic(). A string-producing one-column expression that includes NOW() can therefore be lifted into an ExpressionVirtualColumn; its single-input plan uses SingleStringInputDeferredEvaluationExpressionDimensionSelector and evaluates the expression once per dictionary value instead of per row, changing filtered or grouped results for nondeterministic expressions.

Suggestion: Reject nondeterministic binding analyses here (or force the plain expression path) and add a regression test covering a one-column nondeterministic delegate.

Copy link
Copy Markdown
Member Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

i guess any expression selector that doesn’t run the expression every row has that problem with such an expression; that flag is pretty new and only used by now() which is only a native expression so it currently cannot happen here, and I don't think this is the appropriate place to fix it. Like, i think the expression planner layer should handle it and not allow it to use any of the special path selectors like deferred dictionary encoded evaluation on lookup or caching selectors or similar.

I will look into fixing this problem at that layer in a follow-up PR.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I'm not sure this is really a problem. We shouldn't guarantee that non-deterministic expressions necessarily run on every row, it ties our hands too much in situations like this. I believe it is fine to have the policy that non-deterministic expressions will run one or more times during the query, without promises beyond that.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Understood. Given the intended policy that nondeterministic expressions may run one or more times during a query, the dictionary-deferred evaluation here is acceptable and I withdraw this finding. I rechecked the current delegate and selector path and found no separate actionable issue in this thread.

@gianm gianm left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

A little birdie told me that the following queries would involve regressions or behavior changes on this branch. Please look into them.

SELECT COUNT(*) FROM druid.foo WHERE MV_FILTER_ONLY(dim3, ARRAY['zzz']) IS NULL
SELECT COUNT(*) FROM druid.foo WHERE MV_FILTER_ONLY(dim3, ARRAY['zzz']) >= 'a'
SELECT COUNT(*) FROM druid.foo f INNER JOIN lookup.lookyloo l ON f.dim2 = l.k WHERE MV_FILTER_ONLY(UPPER(l.v), ARRAY['XA']) = 'XA'
SELECT f.dim1, l.v, l2.v FROM druid.foo f INNER JOIN lookup.lookyloo l ON MV_FILTER_ONLY(LOWER(f.dim3), ARRAY['a']) = l.k INNER JOIN lookup.lookyloo l2 ON f.dim2 = l2.k
SELECT f.dim1, l.v, l2.v FROM druid.foo f INNER JOIN lookup.lookyloo l ON f.dim2 = l.k INNER JOIN lookup.lookyloo l2 ON MV_FILTER_ONLY(SUBSTRING(l.v, 2), ARRAY['a']) = l2.k
SELECT MV_FILTER_REGEX(LOWER(dim2), '^a'), COUNT(*) FROM druid.foo GROUP BY 1
SELECT MV_FILTER_ONLY(NVL(dim3, 'x'), ARRAY['x']), COUNT(*) FROM druid.foo GROUP BY 1
SELECT MV_FILTER_ONLY(NVL(dim3, 'x'), ARRAY['x', 'a']), COUNT(*) FROM (SELECT dim3 FROM druid.foo LIMIT 10) GROUP BY 1
SELECT COALESCE(MV_FILTER_ONLY(UPPER(dim3), ARRAY['B']), 'none'), COUNT(*) FROM druid.foo GROUP BY 1
SELECT MV_FILTER_NONE(LOOKUP(dim3, 'lookyloo'), ARRAY['xa']), COUNT(*) FROM druid.foo GROUP BY 1 ORDER BY 2 DESC LIMIT 2

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants