Skip to content

rfc: analyzed lexical search on the shared expression model - #792

Draft
ragnorc wants to merge 6 commits into
mainfrom
rfc/analyzed-lexical-search
Draft

ragnorc wants to merge 6 commits into
mainfrom
rfc/analyzed-lexical-search

Conversation

@ragnorc

@ragnorc ragnorc commented Sep 28, 2026 •

Copy link
Copy Markdown
Contributor

What this is

This RFC makes text search give the same answer whatever the index state. A String property becomes searchable by declaring @analyzed, which fixes its analyzer and scorer in the schema. One lexical query value, terms(text, mode, max_edits), is used two ways: by match_terms(field, terms(…)) as a predicate in match, and by bm25(field, …) as the ranking in order.

It replaces the analyzed lexical search draft in #606. That draft was built on a staged query syntax that the shared expression model replaced. This version is written against main b14c22c5 and uses the shared expression model's rules: predicates live in match, and the leading ranking call lives in order. It adds a function, an argument type and a schema capability, and no clause.

The problems

Text search answers change with physical index state today:

The proposal

  • @analyzed binds one of three immutable analyzer profiles and a default scorer into the accepted schema.
  • match_terms replaces search, match_text and fuzzy. On an @analyzed field the old spellings compile to it; on any other field they keep today's behavior. Both carry a deprecation diagnostic for one release. The mapping uses mode: any, which keeps today's any-term matching (Lance's default query operator is Or).
  • An exact scan applies the field's analyzer and is the correctness baseline. A native index accelerates only where it proves the same matches.
  • One versioned scorer, bm25_v1, for exact and edit-tolerant queries, with statistics over the whole field.
  • Engine v1 is frozen and refuses the new constructs at its door, naming the switch to engine v2.

Evidence

The lexical oracle, fixtures and upstream receipt are kept at ce5a3012 on the search-contracts-evidence-2026-09 branch; the RFC links them by permalink. The four Lance guards it cites are at 1e0bed40 and return with the PRs that depend on them.

Related

Checks: python3 scripts/check-docs.py, bash scripts/check-agents-md.sh, typos.

Update (2026-10-01)

Aligned with RFC 0047 as merged (#791), and every factual claim checked again against main 92ea5449. Two conflicts are resolved:

  • Deprecation release. No query breaks in it: a legacy spelling on a field without @analyzed keeps today's behavior.
  • T27. It requires @analyzed from the removal release onward. Unindexed full-text search is case-sensitive #747 is closed by RFC 0047's refusal, which this RFC's exact scan later lifts.

Also added:

  • the mapping for two-argument fuzzy (two edits);
  • the row change from today's English index analysis to the declared profile, with english_v1 as the closest profile;
  • rebuild-full-text-indexes with the field profile, and an RFC 0043 certificate that records the field's analyzer fingerprint;
  • a SchemaIR feature name instead of a version.

Corrected: the #749 mechanism, the flat scanner's fuzzy behavior, the evidence citations and the named guards.

# Conflicts:
#	docs/rfcs/README.md

@ragnorc ragnorc left a comment

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Recommendation: approve this draft RFC with one non-blocking documentation correction. I found no blocking design defect at c5d33c2931a14ee3fa360836a2ca420cc6fa52c9. This review does not qualify the future implementation for release.

This PR defines how text search should behave. Today, an index can change which words match. Newly appended rows can also behave differently from indexed rows. The proposal puts text analysis in the schema. It gives filtering and ranking one shared lexical query. An exact evaluator defines the answer, and an index may accelerate only equivalent work. Users gain stable matching across index states. They also accept analyzer migration, different scores, and potentially expensive scans.

Contract and code evidence

  • The observable contract is sound: one accepted snapshot, one field analyzer, complete membership, and typed failure when the engine cannot finish within its budget. Index coverage changes cost. It must not silently change results. The proposal fits the same engine on local files, S3, and Azure. It adds no separate server correctness path.
  • The stated upstream defects match the pin. Lance 11 query execution uses a bare tokenizer for nonzero fuzziness. Expansion stops at a shared term budget. Raising that budget cannot prove completeness.
  • The migration maps ordinary search to mode: any and keeps the two-edit default for fuzzy(f,q). This matches OmniGraph's current calls and Lance's default Operator::Or.
  • The profile change requires real builder and certificate changes. The builder uses InvertedIndexParams::default(). The certificate records one engine-wide generation. The RFC correctly requires a field fingerprint and explicit rebuilds. NFC remains an explicit qualification gate.

Tradeoffs and liability

The likely workload includes repeated text searches combined with selective graph traversals, plus writes that leave partial index coverage. That makes consistent membership valuable. However, a small eligible graph set does not make corpus-wide scoring cheap. Exact fuzzy evaluation adds token comparison work. Exact scoring needs statistics from the full visible field corpus, including rows outside the graph filter. Remote storage can make these reads costly. Until acceleration proves parity, large fields or concurrent searches may hit resource limits. No throughput or latency claim is established here.

The design removes three separate search spellings and their separate evaluators after the migration window. One LexicalQuery serves both filtering and ranking. It also keeps index files and postings under Lance ownership. These choices remove semantic liability at its cause.

The design adds durable analyzer identities, three profiles, a scorer version, a bounded evaluator, and certificate compatibility rules. The transition temporarily retains legacy and analyzed behavior. These are real maintenance costs. The PR adds 494 documentation lines and removes no runtime code. Its eventual liability reduction depends on removing the legacy paths and proving one evaluator across all coverage states.

After five similar changes, immutable profiles and scorer versions could become a large compatibility matrix. Requiring an RFC for each new version is useful. Keep new consumers on the same typed query and evaluator. Do not add a separate vocabulary store or a new execution path for each profile.

One implementation detail deserves an explicit decision: projected score types. The current compiler returns F32 for bm25. The new numeric contract uses float64. Define the public result type and test it with score and tie fixtures when implementing step 3.

Validation and limits

  • Local checks on the exact head passed: documentation validation, AGENTS links, and typos for both changed files.
  • Ten existing compiler tests passed: seven selected by search, three by fuzzy. These check the current language, not the proposed constructs.
  • All 13 archived Decimal fixtures regenerated byte-for-byte. They cover already-analyzed terms. They do not prove NFC integration, bounded execution, or runtime score parity.
  • I inspected the archived Lance guards and checked six relevant local Lance source files against the receipt's SHA-256 hashes. All six matched. I did not rerun those historical Rust guards.
  • CI documentation checks passed on merge commit a1f1a848, which includes this head and base 92ea5449. The workspace test job was skipped. No cloud or performance tests ran locally.

The inline comment corrects the compatibility summary. The detailed migration section already states the right behavior. Keep the listed NFC, scorer, and resource qualification gates before implementation approval.

Comment on lines +375 to +377
- **Results:** rows change only where today's answer depended on index state
or letter case, and scores change on adopting fields; that is the correction
this RFC exists for.

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

[P3] Include analyzer-profile changes in the compatibility summary

This summary promises fewer row changes than the Migration section. Adopting bare @analyzed also removes today's stemming, stop-word removal, ASCII folding, and token-length filter. For example, indexed running matches a query for run today. Under standard_v1, those terms differ even with complete index coverage and identical letter case. Please include analyzer-profile changes here, or refer to the Migration section instead of repeating its rules. This is a non-blocking documentation correction because that section already explains the change.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant