Skip to content

rfc: search plan truth, retrieval algebra, analyzed lexical search, and the query kernel - #606

Closed
ragnorc wants to merge 16 commits into
mainfrom
rfc-0047-search-plan-truth
Closed

ragnorc wants to merge 16 commits into
mainfrom
rfc-0047-search-plan-truth

Conversation

@ragnorc

@ragnorc ragnorc commented Sep 3, 2026 •

Copy link
Copy Markdown
Contributor

What & why

Four draft RFCs for search, restructured from this PR's earlier revision: the design is the same, split so each document can be reviewed and accepted on its own, with nothing in the tree that is a prototype and a rollout that ships one reversible step at a time instead of one cutover.

Review 0047 first. It is 530 lines, self-contained, and Phase A of the rollout; it can be accepted while the other three stay drafts.

  • RFC 0047 — Search plan truth (public, draft; Phase A). T26 refuses a search filter or rank target that is silently dropped today; retrieval becomes a typed field of the lowered plan; bm25/nearest/rrf are projectable; ranked output has a total order; the read envelope gains additive warnings, metrics, retrievals and usage. Changed in this revision: full-text search on a property without @index is a compile-time T27 refusal and a declared-but-unbuilt index is a plan-time FullTextIndexRequired refusal (the RFC 0043 shape), replacing the earlier warning, which left invariant 7 open; and a diagnostics contract (stable code, position or stage, expectation, one fix) for every compile diagnostic and typed read failure.
  • RFC 0048 — Search contracts and retrieval algebra (public, draft; Phases D–F). Explicit rank stages with named sources and metrics, knn/ann, named weighted rrf, per-group selection (limit n of $x per { … }), target identity through fan-out, representation identity, result metadata and coherent follow-up. Decided in this revision: read options are session settings (feat(engine): introduce sessions for settings-aware execution #742); combination is an expression and retrieval is named (fuse(expr) as the general fusion source, rrf_v1 its named policy, explicit-mixing and missing-arm rules); the resource ledger and DataFusion material are requirements handed to the engine version 2 components (rfc: add engine version 2 placeholder #711) together with a rewrite catalogue and an explain contract; an agent-facing surface (stored queries as the typed door, @description on schema declarations, named defaults) and a reader decision rule for descriptors.
  • Analyzed lexical search (public, draft; Phase C). @analyzed with immutable analyzer profiles (NFC plus the pinned Lance tokenizer), one typed terms description consumed by match_terms (in filter) and lexical (in rank), bm25_v1 with field-corpus statistics, an exact scan baseline with index acceleration only after parity. Opt-in per field: graphs that do not adopt it need no rebuild. fuzzy, search and match_text get a one-release deprecation window with a fixed mapping, per the compatibility-surfaces RFC.
  • GQ composition and language evolution (maintainer, draft, not under review yet; Phase G). The shared expression rules, the capability matrix, C1–C4 and the deferred cross-type boundary, plus the kernel the other two now use: eight stage kinds (match, filter, let, rank, group, order, limit, return), exactly two cuts, subqueries as expressions under a reduction, three open sets (pattern items, expression functions, retrieval sources) so extensions add no grammar rule, atomic case-insensitive keywords, the pattern/predicate rule (predicates inside a pattern block only in optional { } and not { }), transparent define, a typed plan input surface, and an empirical instrument (first-try validity, turns to a correct query, repairs by code, over ground truth the graph computes itself) as the tie-break for spellings.

Rollout (RFC 0048 § Rollout): Phases A–G ordered by risk — truth first (0047 plus attribution of the served read floor, #752), the agent door, opt-in representations, kernel stages additive behind a session setting and A/B-measured against the current grammar, a deprecate-then-remove cutover, acceleration behind parity oracles, programmability — each with owner, gate, exit and open items, a risk register, and what is off the critical path. No phase depends on a mandatory format rebuild.

Also in the tree: the upstream-contract receipt (assets/0048-upstream-contract-checkpoint.json, the source-audit record) and one green GQT case pinning the field-corpus statistics decision. Probe scripts and oracle fixtures are cited from the evidence branch, not carried here.

Kept out of the tree, retained on search-contracts-evidence-2026-09 (ce5a3012) and cited from the RFCs: the long-form agent-context copy, the two integration patches and receipts, the agent pilot record, the test-only staged compiler and goldens, the DataFusion selection/composition probes, the selection cost bench, the oracle test and its fixtures, and the Arrow pool probe. None tests a current OmniGraph boundary; several would not survive a rebase.

Backing issue / RFC

  • Is an RFC PR: docs/rfcs/0047-search-plan-truth.md, docs/rfcs/0048-search-contracts.md, docs/rfcs/2026-09-18-analyzed-lexical-search.md, docs/rfcs/2026-09-18-gq-composition-and-language-evolution.md (0047 and 0048 land under the numbers the registry reserved for this PR when the namespace closed).

Checklist

  • Change is focused (RFC documents, their receipts, one GQT case)
  • Tests added/updated for behavior changes (N/A: no behavior changes; the substrate probes are a separate PR)
  • Public docs updated if user-facing surface changed (N/A: drafts)
  • Reviewed against docs/dev/invariants.md — each RFC's Invariants section; no deny-list item hit

Local verification

  • python3 scripts/check-docs.py — Documentation OK (147 Markdown files)
  • bash scripts/check-agents-md.sh — OK
  • typos — clean
  • git diff --check — clean
  • cargo test -p omnigraph-gqt --locked --test gq_logic_tests bm25_statistics_ignore_eligibility_filter — 1 passed on main 8281807b
  • RFC 0047's rollout now lists Phase A as nine PRs against the current tree with a stored-query deployment gate; the composition RFC records a first in-context measurement on today's grammar (45/45 correct, 37/45 first-try valid, one error class)
  • cargo test --workspace — not run: no Rust files changed

Notes for reviewers

@chatgpt-codex-connector

chatgpt-codex-connector Bot commented Sep 3, 2026 •

Copy link
Copy Markdown

Codex Review Summary

This comment shows the latest Codex review activity on this pull request.

Review Status Commit Review trigger
📝 Code Review ✅ Completed 2026-09-03T11:47:58.721607Z 5b7ce5a PR opened
ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review" or "@codex security review".

Codex reacts with 👀 while any review is running, comments if it has suggestions, and reacts with 👍 once all reviews finish with no findings.

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: 5b7ce5a344

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

Comment thread docs/rfcs/0047-search-plan-truth.md Outdated
Comment on lines +184 to +187
- **Physical acceleration is derived (7):** preserved — the unindexed-column
condition warns, it does not fail; index absence changes cost and (on the
flat fallback) analysis behavior, which is exactly what the warning makes
visible. Recall reporting is contractual, not plan-derived.

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P1 Badge Prevent missing FTS indexes from changing search answers

When the requested column has no FTS index, this design explicitly permits a case-sensitive fallback whose results differ from indexed execution and only adds a warning. The same query can therefore gain or lose rows after index reconciliation even though an index is derived state; making that discrepancy visible does not restore logical correctness. Use analyzer-equivalent fallback behavior or fail closed instead.

AGENTS.md reference: AGENTS.md:L99-L100

Useful? React with 👍 / 👎.

Comment thread docs/rfcs/0047-search-plan-truth.md Outdated
Comment on lines +171 to +174
- **Coverage.** Ready/pending counts reuse the scan's own structured
predicate through a sealed, streaming count on the storage boundary — no
SQL strings, no retained batches, computed only for `@embed`-backed vector
retrievals.

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P1 Badge Bound coverage work independently of table size

For an indexed nearest query with a small limit and no selective filter, computing exact ready/pending counts through a streaming predicate still scans the entire source/vector population on every request. Streaming bounds retained memory, but not the O(table-size) I/O and latency added to an otherwise sublinear ANN read; coverage must use bounded metadata, be opt-in, or otherwise avoid an unconditional full-population count.

AGENTS.md reference: AGENTS.md:L104-L105

Useful? React with 👍 / 👎.

Comment thread docs/rfcs/0047-search-plan-truth.md Outdated
Comment on lines +24 to +26
1. `fuzzy()` is retired with a stable `T25` compile diagnostic — it provably
never matched under the supported tokenizer, so every use was a confident
empty answer.

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P1 Badge Prove fuzzy is universally inert before retiring it

The cited characterization in crates/omnigraph/tests/search.rs:1764-1787 exercises only one capitalized, stem-sensitive typo (Introductio), which cannot establish that every fuzzy() invocation returns empty. Fuzzy matching includes zero-edit matches, so normalized terms such as lower-case deep can still match indexed terms; rejecting every use with T25 would break working stored queries despite the compatibility section calling all affected usage provably broken. Add representative exact/stem/case/max-edit evidence or narrow the retirement to the actually inert shape.

AGENTS.md reference: AGENTS.md:L151-L153

Useful? React with 👍 / 👎.

Comment thread docs/rfcs/0047-search-plan-truth.md Outdated
Comment on lines +236 to +240
A complete prototype exists (closed PR #595, branch
`search-contracts-p0-p1`, retained as evidence per the closure note): eleven
staged commits, canonical workspace graph green (2,860 tests), both Clippy
gates, OpenAPI regenerated, vocabulary-guard inventory classified. Test
owners extended, not forked: compiler typecheck/lowering suites (T25, T26,

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Badge Record the exact Lance surfaces reviewed

This Lance-dependent RFC reports only that an unspecified impact analysis and source validation occurred; it does not identify the complete upstream index, FTS, tokenizer, vector, or DataFusion pages reviewed. The RFC process requires the exact version and surveyed surfaces so acceptance can verify that the fallback, tie, and index-lifecycle assumptions were checked against the pinned substrate rather than an unavailable prototype branch.

AGENTS.md reference: AGENTS.md:L13-L17

Useful? React with 👍 / 👎.

@ragnorc ragnorc changed the title rfc: add RFC 0047, search plan truth rfc: search contracts program — RFC 0047 (search plan truth) and RFC 0048 (search contracts and retrieval algebra) Sep 3, 2026
@ragnorc ragnorc changed the title rfc: search contracts program — RFC 0047 (search plan truth) and RFC 0048 (search contracts and retrieval algebra) rfc: search contracts and retrieval algebra Sep 3, 2026
aaltshuler added a commit that referenced this pull request Sep 3, 2026
0047 and 0048 are allocated by PR #606, so the RFC becomes 0049 and the
registry's next number 0050. From review: the ledger restore seam is
withdrawn, since its real use is a coherent restore point where the ledger
and the graphs come back together; readiness reports counts only and the
graph ids stay behind the authenticated GET /graphs; the shutdown watchdog
is an operating-system thread armed by a listener installed before graphs
open; the authority enum gains `unlocked`; and the compatibility section
states the Rust-level breaks instead of calling the change additive.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015cD8PEeUfrzpqYuU1jaZBq
@aaltshuler aaltshuler mentioned this pull request Sep 3, 2026
5 of 7 tasks
aaltshuler added a commit that referenced this pull request Sep 3, 2026
0047 and 0048 are allocated by PR #606, so the RFC becomes 0049 and the
registry's next number 0050. From review: the ledger restore seam is
withdrawn, since its real use is a coherent restore point where the ledger
and the graphs come back together; readiness reports counts only and the
graph ids stay behind the authenticated GET /graphs; the shutdown watchdog
is an operating-system thread armed by a listener installed before graphs
open; the authority enum gains `unlocked`; and the compatibility section
states the Rust-level breaks instead of calling the change additive.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015cD8PEeUfrzpqYuU1jaZBq
aaltshuler added a commit that referenced this pull request Sep 3, 2026
rfc: accept RFC 0049

Maintainer decision recorded in the decision log; the registry row moves
to accepted. Implementation status advances with the two implementation
PRs.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015cD8PEeUfrzpqYuU1jaZBq

rfc: renumber to 0049, drop the ledger restore seam, record the review

0047 and 0048 are allocated by PR #606, so the RFC becomes 0049 and the
registry's next number 0050. From review: the ledger restore seam is
withdrawn, since its real use is a coherent restore point where the ledger
and the graphs come back together; readiness reports counts only and the
graph ids stay behind the authenticated GET /graphs; the shutdown watchdog
is an operating-system thread armed by a listener installed before graphs
open; the authority enum gains `unlocked`; and the compatibility section
states the Rust-level breaks instead of calling the change additive.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015cD8PEeUfrzpqYuU1jaZBq

rfc: add RFC 0048, control-plane seams

Four small, independently shippable contracts an external control plane
needs from the cluster crate and the server: observe-only reads
(`plan --observe`, `cluster observe`), a ledger restore
(`cluster state restore`), a readiness witness (`GET /readyz`), and a
bounded shutdown (`--shutdown-grace-seconds`). Nothing changes a storage
format, the ledger, the lock, or the recovery protocol; RFC 0034 and 0035
stay independent. The next number becomes 0049: 0047 is held by an
out-of-tree draft.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015cD8PEeUfrzpqYuU1jaZBq
@azimafroozeh azimafroozeh mentioned this pull request Sep 4, 2026
5 tasks done
@ragnorc

ragnorc commented Sep 8, 2026

Copy link
Copy Markdown
Contributor Author

Azim, the latest RFC 0048 iteration makes bounded retrieval compose explicitly with ordinary graph queries. This covers the staged retrieval redesign and the follow-up migration, implementation-phase, and default changes through 5746a072 (revision diff).

  1. Retrieval is an explicit stage in .gq. Predicates establish eligibility; rank stages select targets; later graph operations work on those results. Named lexical/vector sources and weighted fusion carry their metrics through the plan. Source windows, per-group selection, and final output limits have separate meanings. Filtering before a candidate cut and filtering afterward are deliberately different operations. Ranked relations are internal; the public surface extends the existing clause language.

  2. The agent workflow uses ordinary graph identity and stored queries. Compact discovery leads to selective source reads, graph expansion, and exact verification at a coherent snapshot. Document and Passage remain application-defined types; there is no built-in EvidenceReference or separate retrieval-profile registry. Metadata distinguishes completion, representation coverage, selection, and metric origin. A relevance score is not answer confidence, and an empty bounded search is not proof of absence.

  3. The initial release now has a concrete boundary. It includes analyzed matching, exact and fuzzy lexical ranking using the same typed lexical query, exact knn, approximate ann, named fusion, graph-defined scope, per-group selection, and coherent source reads. Learned reranking, sparse/multivector representations, ranked pagination, and snippets remain explicit extensions. Fuzzy ranked retrieval and whole-query resource bounds are required for completion.

  4. Breaking changes and rollout are explicit. One pre-stable cutover removes the legacy lexical spellings, nearest, retrieval inside order, and positional RRF, without a compatibility release. The new migration section separates query rewrites from additive capabilities and explains the accepted-schema export/init/load rebuild, including its effects on history, branches, and indexes. Seven implementation phases cover contract resolution, shared foundations, exact retrieval, graph composition, agent-facing reads, physical qualification/evaluation, and the coordinated migration.

  5. Schema defaults are useful and reproducible. Bare @analyzed expands to standard_v1 plus BM25-family capability; scorer="none" opts out of ranking. Embedding fields can inherit a schema-owned qualified recipe. Source and dimensions remain required; distance can inherit from that recipe, while raw vectors require explicit geometry. All choices resolve into persisted field semantics, appear in schema plan, and survive export/reapplication. Runtime provider defaults cannot silently change an accepted field. The proposed analyzer profiles now apply NFC before tokenization; stored strings and exact predicates retain their original behavior.

  6. Validation claims are separated from acceptance gates. The audit records pinned Lance/DataFusion mechanisms, 442 passing current-code tests, and the 73,008-comparison tokenizer/edit-distance probe. Those checks do not qualify the new staged operators, schema defaults, NFC pipeline, or end-to-end budgets. Exact native paths need semantic parity with the exact baseline; ANN needs recall and effort qualification. All paths share resource accounting and fail explicitly when their contract cannot be completed.

The main review points are stage/metric grammar and identity through graph fan-out, the schema-default declaration and migration rules, and the remaining scoring, encoding-identity, and NFC/index qualification gates. Both RFCs remain drafts with implementation not started. This revision changes documentation only; documentation checks passed.

@aaltshuler
aaltshuler marked this pull request as draft September 13, 2026 11:58
ragnorc added a commit that referenced this pull request Sep 18, 2026
Takes the two RFC documents, the upstream-contract receipt, the Arrow pool
probe, the lexical scoring oracle fixtures (moved under assets/) and the one
green GQT case (ported to the corpus's `--- runner` header) from PR #606 at
ce5a301, on top of main's closed RFC number namespace.

The agent-context copy of the RFC, the archived integration patches, the
agent pilot record, the test-only staged compiler, the DataFusion
composition probes, the selection bench and the oracle test stay on the
`rfc-0047-search-plan-truth` branch as evidence; none of them tests a
current OmniGraph boundary.
@ragnorc
ragnorc force-pushed the rfc-0047-search-plan-truth branch from ce5a301 to ede8004 Compare September 18, 2026 11:21
@ragnorc ragnorc changed the title rfc: search contracts and retrieval algebra rfc: search plan truth, retrieval algebra, analyzed lexical search, and the query kernel Sep 18, 2026
ragnorc added a commit that referenced this pull request Sep 18, 2026
Takes the two RFC documents, the upstream-contract receipt, the Arrow pool
probe, the lexical scoring oracle fixtures (moved under assets/) and the one
green GQT case (ported to the corpus's `--- runner` header) from PR #606 at
ce5a301, on top of main's closed RFC number namespace.

The agent-context copy of the RFC, the archived integration patches, the
agent pilot record, the test-only staged compiler, the DataFusion
composition probes, the selection bench and the oracle test stay on the
`rfc-0047-search-plan-truth` branch as evidence; none of them tests a
current OmniGraph boundary.
@ragnorc
ragnorc force-pushed the rfc-0047-search-plan-truth branch from 078be9c to 924d7b3 Compare September 18, 2026 13:21
ragnorc added a commit that referenced this pull request Sep 19, 2026
…ion-scoped migration, no whole-graph rebuild, filter in Phase C, one nullable-metric rule)
Takes the two RFC documents, the upstream-contract receipt, the Arrow pool
probe, the lexical scoring oracle fixtures (moved under assets/) and the one
green GQT case (ported to the corpus's `--- runner` header) from PR #606 at
ce5a301, on top of main's closed RFC number namespace.

The agent-context copy of the RFC, the archived integration patches, the
agent pilot record, the test-only staged compiler, the DataFusion
composition probes, the selection bench and the oracle test stay on the
`rfc-0047-search-plan-truth` branch as evidence; none of them tests a
current OmniGraph boundary.
RFC 0048 keeps the staged retrieval algebra. The lexical contract moves to
2026-09-18-analyzed-lexical-search (`@analyzed`, analyzer profiles, `terms`,
`match_terms`, `bm25_v1`, the exact scan baseline, the regression cases);
the grammar rules, capability matrix, C1–C4 and cross-type discovery move
to 2026-09-18-gq-composition-and-language-evolution. Sections move verbatim;
links are repointed and the evidence branch is cited by commit.

Decisions recorded on the way:

- The legacy spellings (`fuzzy`, `search`, `match_text`, `nearest`,
  retrieval in `order`, positional `rrf`) get a one-release deprecation
  window with a fixed mapping instead of a hard removal, following the
  compatibility-surfaces RFC; stored queries are re-parsed at boot.
- The read options (`coverage`, `require_replay`, limits) are session
  settings after #742 landed `Session`, not a new request field.
- The resource ledger and the DataFusion lowering sections are marked as
  requirements handed to the engine version 2 components named by RFC 0067
  (PR #711); RFC 0048 no longer decides a planner or an admission mechanism.
- RFC 0047 replaces the `full_text_search_unindexed` warning with a
  compile-time `T27` refusal of undeclared full-text targets and a plan-time
  `FullTextIndexRequired` refusal of unbuilt indexes (the RFC 0043 fence
  shape); a warning left invariant 7 open.
Adds the kernel section to the GQ composition RFC: eight stage kinds
(match, filter, let, rank, group, order, limit, return), exactly two cuts (a
retriever's `candidates:` window and `limit`), subqueries as correlated
expressions under a reduction (collect, one, exists, count), three open sets
the typechecker resolves (pattern items, expression functions, retrieval
sources) so extensions add no grammar rule, and atomic-positional keywords
(the RFC 0056 rule). It collapses RFC 0048's `select`, `take`, `score`,
`collect` and `optional` spellings, moves scalar predicates out of `match`
(the mechanism behind the dropped-predicate class RFC 0047's T26 refuses),
and gives sources a `ties:` option so trailing order keys no longer decide a
cut. Cypher/GQL constructs and the agent retrieval needs are mapped onto it;
the RFC examples are shown in kernel form.

RFC 0048 cross-references the proposal as a pending decision; its own
semantics and examples are unchanged.
The caller is a general agent trained with RL elsewhere, using the language
in context; nothing here assumes training on it or proximity to a prior
language.

- RFC 0047: a diagnostics contract for every compile diagnostic and typed
  read failure — stable code, position or stage, expectation, one fix;
  unknown names enumerate their set. No catalogue of foreign idioms.
- RFC 0048: an agent-facing surface (stored queries as the typed door,
  `@description` on schema declarations surfaced by `schema show`, named
  defaults); a reader decision rule putting completion, coverage, selection
  and usage in the response body, with usage as contract; the in-context
  competence metric in the mixed-workload qualification.
- Composition RFC: the in-context consumer as the kernel's design target
  (card-sized, regular, one-turn repair), case-insensitive atomic keywords,
  and an empirical instrument — first-try validity, turns to a correct
  query, repairs by code, over ground truth the graph computes itself — as
  the tie-break for spellings and the gate for new stages.
…logue

RFC 0048: combination is an expression, retrieval is named. `fuse(expr,
candidates:)` is the general fusion source, evaluated over the union of
named arms; `rrf` is its named, qualified policy `rrf_v1`. Explicit mixing
only, missing arms explicit (coalesce or nulls placement), structural
identity for inline formulas, models enter as budgeted sources, control
flow stays outside the query. A rewrite catalogue (always-legal rewrites,
barriers, cost-based choices with visible statistics and differential
oracles, the `ann`-only approximation boundary, in-attempt adaptivity, plan
cache) and an `explain` contract are handed to the engine version 2
planner.

Composition RFC: programmability through transparent `define`, a typed plan
input surface and multi-statement requests at one snapshot; procedural
loops, opaque functions and free inference calls excluded.
…over

RFC 0048's rollout no longer promises a single pre-stable change with a
mandatory format rebuild. Phases A–G ship one reversible step at a time:
truth first (RFC 0047 plus session-scoped caching), the agent door (stored
queries, `@description`, the in-context instrument and its baseline),
representations opt-in per field (no rebuild for non-adopters), the kernel
stages additive behind a session setting and A/B-measured against the old
grammar, a deprecate-then-remove cutover, acceleration behind parity
oracles, and programmability. Each phase names its owner RFC, gate, exit and
open items; a risk register says where each assumption is tested and what
happens if it is wrong. The lexical RFC becomes Phase C, RFC 0047 Phase A,
the composition RFC Phase G; the Phase 0 design gates keep their name.
…exical RFC

RFC 0048's per-group selection is `limit n of $x per { … }` on the active
order; `let` carries scoring features; `filter` carries predicates; the
kernel note records the adoption. The composition RFC's C1–C4 are rewritten
to `filter`, `order` + `limit`, `let` scoring features and `sub()` under
`one`/`collect`; its capability rows and operator table follow; the
duplicated C4 in the kernel section is dropped in favour of the example.
The lexical RFC moves `match_terms` from the `match` block to a `filter`
stage and its migration mapping says so. No semantic decision changed.
…ion-scoped migration, no whole-graph rebuild, filter in Phase C, one nullable-metric rule)
@ragnorc
ragnorc force-pushed the rfc-0047-search-plan-truth branch from 30d56e4 to ad37e5c Compare September 25, 2026 08:20
ragnorc added a commit that referenced this pull request Sep 25, 2026
Six lance_surface_guards probes and one CLI test from PR #606, applied on
main; none of them depends on the RFCs being accepted:

- fts_statistics_scope_can_reverse_ranking: eligibility and the BM25
  scoring corpus are independent (an external row mask keeps the scores an
  index rebuild over the subset would reverse).
- nfc_preprocessing_requires_an_explicit_bounded_integration: the public
  tokenizer can be reused after NFC, but no index-builder option adds it.
- native_fuzzy_bm25_rewards_a_rare_expansion and
  native_bm25_idf_can_round_a_common_term_to_zero: why native BM25 is not
  the proposed scorer (rare expansion outweighs the exact term; f32 IDF of
  a common term rounds to zero at N = 2^24).
- lance_provider_scan_payloads_need_accounting_beyond_the_session_pool and
  native_scheduler_drop_distinguishes_queued_and_dispatched_reads: the
  session pool does not bound decoded Lance batches, and dropping a result
  future cancels a queued read but not a dispatched one.
- cli_data: a recorded model label does not freeze a provider's output at a
  snapshot; explicit vectors make no provider call and a mismatched label
  is refused.

The IDF reference values use f64::ln_1p instead of the pinned libm
dev-dependency the PR carried; these probes need a positive reference, not
bit identity.
@ragnorc

ragnorc commented Sep 28, 2026

Copy link
Copy Markdown
Contributor Author

Closing this in favour of three fresh PRs, each written from scratch against main b14c22c5:

Why start over rather than keep editing here: since this PR opened, engine v2 and the shared expression model landed on main. They change where each fix belongs, and several assumptions in these drafts are no longer true (for example the staged rank … yield syntax, and the claim that search() on a traversal destination is still wrong, which #760 fixed on engine v2). Rewriting from validated facts avoids carrying outdated statements forward.

What changed:

  • RFC 0047 now lands on engine v2 and the shared compiler. Engine v1 is frozen and refuses the shapes it would answer wrongly. A ranked binding becomes the root of its component instead of being refused, rrf() cuts rows, and a search-ordered aggregate applies its remaining keys. It is back to draft until one question for the engine owner is answered.
  • Analyzed lexical search keeps the analyzer profiles, terms, the matching laws and bm25_v1, but match_terms is a search call in match and ranking stays bm25(…) in order. The deprecation mapping now uses mode: any, which keeps today's any-term matching.
  • RFC 0048 withdraws the staged syntax and keeps its semantics as laws over the planner's retrieval nodes, with named options on the existing ranking calls.
  • GQ composition and language evolution is withdrawn. Its agent design target and the in-context competence measurement move into RFC 0048.

RFC numbers 0047 and 0048 stay reserved and move to #791 and #793. The evidence this PR produced (the lexical oracle, fixtures, probes and receipts) stays at ce5a3012 on the search-contracts-evidence-2026-09 branch, and the new RFCs link to it by permalink.

@ragnorc ragnorc closed this Sep 28, 2026
@ragnorc
ragnorc deleted the rfc-0047-search-plan-truth branch September 28, 2026 13:47
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants