Repository navigation
Conversation
# Conflicts: # docs/rfcs/README.md
ragnorc
left a comment
There was a problem hiding this comment.
Recommendation: approve this draft RFC with one non-blocking documentation correction. I found no blocking design defect at c5d33c2931a14ee3fa360836a2ca420cc6fa52c9. This review does not qualify the future implementation for release.
This PR defines how text search should behave. Today, an index can change which words match. Newly appended rows can also behave differently from indexed rows. The proposal puts text analysis in the schema. It gives filtering and ranking one shared lexical query. An exact evaluator defines the answer, and an index may accelerate only equivalent work. Users gain stable matching across index states. They also accept analyzer migration, different scores, and potentially expensive scans.
Contract and code evidence
- The observable contract is sound: one accepted snapshot, one field analyzer, complete membership, and typed failure when the engine cannot finish within its budget. Index coverage changes cost. It must not silently change results. The proposal fits the same engine on local files, S3, and Azure. It adds no separate server correctness path.
- The stated upstream defects match the pin. Lance 11 query execution uses a bare tokenizer for nonzero fuzziness. Expansion stops at a shared term budget. Raising that budget cannot prove completeness.
- The migration maps ordinary search to
mode: anyand keeps the two-edit default forfuzzy(f,q). This matches OmniGraph's current calls and Lance's defaultOperator::Or. - The profile change requires real builder and certificate changes. The builder uses
InvertedIndexParams::default(). The certificate records one engine-wide generation. The RFC correctly requires a field fingerprint and explicit rebuilds. NFC remains an explicit qualification gate.
Tradeoffs and liability
The likely workload includes repeated text searches combined with selective graph traversals, plus writes that leave partial index coverage. That makes consistent membership valuable. However, a small eligible graph set does not make corpus-wide scoring cheap. Exact fuzzy evaluation adds token comparison work. Exact scoring needs statistics from the full visible field corpus, including rows outside the graph filter. Remote storage can make these reads costly. Until acceleration proves parity, large fields or concurrent searches may hit resource limits. No throughput or latency claim is established here.
The design removes three separate search spellings and their separate evaluators after the migration window. One LexicalQuery serves both filtering and ranking. It also keeps index files and postings under Lance ownership. These choices remove semantic liability at its cause.
The design adds durable analyzer identities, three profiles, a scorer version, a bounded evaluator, and certificate compatibility rules. The transition temporarily retains legacy and analyzed behavior. These are real maintenance costs. The PR adds 494 documentation lines and removes no runtime code. Its eventual liability reduction depends on removing the legacy paths and proving one evaluator across all coverage states.
After five similar changes, immutable profiles and scorer versions could become a large compatibility matrix. Requiring an RFC for each new version is useful. Keep new consumers on the same typed query and evaluator. Do not add a separate vocabulary store or a new execution path for each profile.
One implementation detail deserves an explicit decision: projected score types. The current compiler returns F32 for bm25. The new numeric contract uses float64. Define the public result type and test it with score and tie fixtures when implementing step 3.
Validation and limits
- Local checks on the exact head passed: documentation validation, AGENTS links, and typos for both changed files.
- Ten existing compiler tests passed: seven selected by
search, three byfuzzy. These check the current language, not the proposed constructs. - All 13 archived Decimal fixtures regenerated byte-for-byte. They cover already-analyzed terms. They do not prove NFC integration, bounded execution, or runtime score parity.
- I inspected the archived Lance guards and checked six relevant local Lance source files against the receipt's SHA-256 hashes. All six matched. I did not rerun those historical Rust guards.
- CI documentation checks passed on merge commit
a1f1a848, which includes this head and base92ea5449. The workspace test job was skipped. No cloud or performance tests ran locally.
The inline comment corrects the compatibility summary. The detailed migration section already states the right behavior. Keep the listed NFC, scorer, and resource qualification gates before implementation approval.
| - **Results:** rows change only where today's answer depended on index state | ||
| or letter case, and scores change on adopting fields; that is the correction | ||
| this RFC exists for. |
There was a problem hiding this comment.
[P3] Include analyzer-profile changes in the compatibility summary
This summary promises fewer row changes than the Migration section. Adopting bare @analyzed also removes today's stemming, stop-word removal, ASCII folding, and token-length filter. For example, indexed running matches a query for run today. Under standard_v1, those terms differ even with complete index coverage and identical letter case. Please include analyzer-profile changes here, or refer to the Migration section instead of repeating its rules. This is a non-blocking documentation correction because that section already explains the change.
What this is
This RFC makes text search give the same answer whatever the index state. A String property becomes searchable by declaring
@analyzed, which fixes its analyzer and scorer in the schema. One lexical query value,terms(text, mode, max_edits), is used two ways: bymatch_terms(field, terms(…))as a predicate inmatch, and bybm25(field, …)as the ranking inorder.It replaces the analyzed lexical search draft in #606. That draft was built on a staged query syntax that the shared expression model replaced. This version is written against
mainb14c22c5and uses the shared expression model's rules: predicates live inmatch, and the leading ranking call lives inorder. It adds a function, an argument type and a schema capability, and no clause.The problems
Text search answers change with physical index state today:
searchandmatch_textrun a bare, case-sensitive scan:"deep"and"Deep"match different rows (Unindexed full-text search is case-sensitive #747).fuzzyspends its edit budget on letter case, because the query is not lowercased while the index is (fuzzy() bypasses the index analyzer at a nonzero edit budget #748).betomissesbetaandrunningmisses the stored stemrun(Index state changes which rows a text query matches #749).The proposal
@analyzedbinds one of three immutable analyzer profiles and a default scorer into the accepted schema.match_termsreplacessearch,match_textandfuzzy. On an@analyzedfield the old spellings compile to it; on any other field they keep today's behavior. Both carry a deprecation diagnostic for one release. The mapping usesmode: any, which keeps today's any-term matching (Lance's default query operator isOr).bm25_v1, for exact and edit-tolerant queries, with statistics over the whole field.Evidence
The lexical oracle, fixtures and upstream receipt are kept at
ce5a3012on thesearch-contracts-evidence-2026-09branch; the RFC links them by permalink. The four Lance guards it cites are at1e0bed40and return with the PRs that depend on them.Related
Checks:
python3 scripts/check-docs.py,bash scripts/check-agents-md.sh,typos.Update (2026-10-01)
Aligned with RFC 0047 as merged (#791), and every factual claim checked again against
main92ea5449. Two conflicts are resolved:@analyzedkeeps today's behavior.T27. It requires@analyzedfrom the removal release onward. Unindexed full-text search is case-sensitive #747 is closed by RFC 0047's refusal, which this RFC's exact scan later lifts.Also added:
fuzzy(two edits);english_v1as the closest profile;rebuild-full-text-indexeswith the field profile, and an RFC 0043 certificate that records the field's analyzer fingerprint;Corrected: the #749 mechanism, the flat scanner's fuzzy behavior, the evidence citations and the named guards.