diff --git a/docs/branch-review-records/cc149b75e900816f35298154ab35a6e553ac49a88cacbd0b7ea4de568250259a.record.md b/docs/branch-review-records/cc149b75e900816f35298154ab35a6e553ac49a88cacbd0b7ea4de568250259a.record.md new file mode 100644 index 0000000000..580311669b --- /dev/null +++ b/docs/branch-review-records/cc149b75e900816f35298154ab35a6e553ac49a88cacbd0b7ea4de568250259a.record.md @@ -0,0 +1 @@ +| 2026-08-18 | claude/s3-follow-up-suggestions-95e160 | b0544fcafaf9effad83b3a379e580ccf68c921fa | packet S3 / A4: menu-derived, evidence-gated follow-up chips in src/lib/answer-follow-up.ts + focused unit and DOM proof | Approved - deterministic composition only; candidates come from the S2 related-information menu, gated on retrieved evidence, suppressed when the answer or an emitted section of that kind already covers them; no src/lib/rag edit, no ClinicalDashboard change, no new module/field/render block; generation prompt untouched | answer-follow-up 27/27; answer-follow-up-chips DOM 6/6; answer-composition 9/9; lint; typecheck; build compiled 2.6min + client bundle secret check; eval:rag:offline 26 suites/623 tests (36 golden cases); eval:rag:adversarial:offline 25/25; check:medication-interactions; check:medication-lexicon-report; full unit suite 7079 passed / 8 failed in 6 unrelated files - session-start-hook reproduced at merge base e1749bf8d, the other five pass in isolation on this branch (Windows parallel-load timeouts) | diff --git a/docs/rag-behaviour/behaviour-map.md b/docs/rag-behaviour/behaviour-map.md index 96ff7fbdc5..c5d181964a 100644 --- a/docs/rag-behaviour/behaviour-map.md +++ b/docs/rag-behaviour/behaviour-map.md @@ -134,3 +134,9 @@ answerShape}` where `answerShape` is provider-safe counts/lengths only — never ladder apply to menu sections unchanged. Contract pins: `tests/answer-composition.test.ts` (all 48 class×intent cells), `tests/rag-answer-composition-prompt.test.ts`, `tests/rag-answer-fallback.test.ts` (menu line in the real prompt input). +- Downstream consumer (packet S3, 2026-08-18): `src/lib/answer-follow-up.ts` reads the same menu to + compose the answer surface's follow-up chips — menu-derived candidates, gated on the retrieved + evidence the client actually received, and suppressed when the answer body or an emitted section of + that kind already covers them. It only READS the menu and changes no retrieval, ranking or prompt + surface. Its templates are index-aligned with the menu items, asserted over all 48 cells in + `tests/answer-follow-up.test.ts`, so a menu edit fails offline instead of silently dropping a chip. diff --git a/docs/rag-improvement/HANDOVER.md b/docs/rag-improvement/HANDOVER.md index 430d85116d..b9093fc0c6 100644 --- a/docs/rag-improvement/HANDOVER.md +++ b/docs/rag-improvement/HANDOVER.md @@ -51,9 +51,11 @@ generation-quality verdict on fallback`), merged 2026-08-13 — structured #2065 (condition-first for/in binding) regressed `agitation-im-po-route-short-terms` live and was reverted by PR #2088; the confirmation run 32100681177 on `4ea310e48` is green (recall 1.0/1.0, zero rr regressions, answer gate 44/44) and is the baseline half of the S2 canary - pair. **Do not reintroduce condition-first for/in binding.** S2 (A2 + A3) opened 2026-08-18: - `src/lib/rag/answer-composition.ts`, prompt `clinical-rag-answer-v19`, `answerSections.maxItems` - 6, adversarial baseline re-captured for v19 (`baseline-record.md` §4). + pair. **Do not reintroduce condition-first for/in binding.** S2 (A2 + A3) merged 2026-08-18 as + squash `dda4956ff`: `src/lib/rag/answer-composition.ts`, prompt `clinical-rag-answer-v19`, + `answerSections.maxItems` 6, adversarial baseline re-captured for v19 (`baseline-record.md` §4). Its canary + pair 32100681177 -> 32111839806 is green, and `eval:answer-quality` was neutral within + nondeterminism (owner blinded read pending). **Track A is complete with S3 (A4).** - **Owner decisions 2026-08-17:** (1) **R1 before S2** — A2/A3 add answer length, and length under the still-unbudgeted strong retry pushes more dosing queries into `provider_timeout`, not fewer; (2) **governance Option B** for the document-summary `similarity: 1` question @@ -74,25 +76,25 @@ generation-quality verdict on fallback`), merged 2026-08-13 — structured ## 2. Status table — update in every programme PR -| Packet | Scope | Branch | PR | State | Canary / evidence refs | -| ---------- | ----------------------------------------------------------------------------------------------------------------------------------------------------- | ------------------------------------------------ | --------------------- | ------------------------------------------------------------------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | -| Guide | Programme guide | `claude/rag-plan-review-guide-vhrls9` | #1895 | Merged 2026-08-13 | docs-only | -| Handover | Multi-session handover + coordination | `claude/rag-plan-review-guide-vhrls9` | #1908 / #2024 | Merged 2026-08-13; coordination layer PR #2024 | docs-only | -| S0 | A1 phase 1: structured fallback diagnostics | `claude/lithium-generation-quality-debug-ji1vce` | #1899 | Merged 2026-08-13 | offline 93/93 focused | -| S1 | A1 phase 2: rung-1 verification-faithfulness fixes | `claude/s1-rag-mitigation-231-86c182` | #2022 | Merged 2026-08-17 (squash `2bd146eed`, landed by content) | 8 pre-fix + 5 post-fix live probes 2026-08-17; offline 583/583; canary pair run 31964560921 (baseline `8f8d111ab`) -> run 32025082010 (`2bd146eed`): recall 1.0/1.0, zero per-case rr regressions, answer gate 44/44 (denominator reconciled by S5; see baseline-record §3); rung-2 measurement in `docs/audit/live-drift-forensics-2026-08.md` §5 | -| S1b | A1 rung 3 (R1): pre-deadline strong routing for dosing class | `claude/s1b-rag-dosing-routing-6u1mik` | #2035 | Merged 2026-08-17 (PR #2035, merge `92f7618`) | canary pair pending: baseline run 32025082010 (`2bd146eed`) -> post-merge dispatch (owner-approved); offline 586/586 + verify:pr-local heavy scope green | -| S1c | A1 residuals R2 + R3: claim-support strictness | `claude/s1c-residuals-r2-r3-4pb1at` | #2052 | Merged 2026-08-17 (merge `b8e774bcd`; follow-up #2063 kept; follow-up #2065 reverted by PR #2088 after canary regression) | canary pair: baseline run 32049952885 -> post run 32052479537 (`084f63799`): recall 1.0/1.0, zero per-case rr regressions, answer gate 44/44 | -| S1d | A1 final-gate gap recovery: hedged cited low-confidence fast answers must recover extractively, not collapse to a citation-free `provider_source_gap` | `claude/s1d-final-gate-gap-recovery-dxgrn2` | #2054 | Merged 2026-08-17 (merge `0bbd64fbc`); landed by content; canary pair green after the #2065 revert (PR #2088) | canary pair 32052479537 -> 32100681177 (`4ea310e48`) green: recall 1.0/1.0, zero per-case rr regressions, answer gate 44/44. Interim post run 32097916649 (`9904fbda8`) was RED on `agitation-im-po-route-short-terms` — bisected live to PR #2065 (S1c follow-up condition-first regex), not S1d; reverted by PR #2088; the confirmation run is 32100681177 | -| G1 | Governance: provenance tag for document-summary rows (Option B) | `claude/g1-rag-document-context-qn9ubx` | #2053 | Merged 2026-08-17 (merge `125e98526`); rows #J912J9 / #0MSNT8 closed at reconcile | no canary (no behaviour change); `document_context` tag at `types.ts` + `rag-row-contracts.ts`, deriveConfidence pinned | -| S2 | A2 + A3: composition menu + moderate length | `claude/s2-rag-composition-7330b0` | #2097 | PR open 2026-08-18 (A2 and A3 together; prompt `clinical-rag-answer-v19`, schema v4, `answerSections.maxItems` 6); owner merges | offline: composition 9/9 + prompt pins 5/5, `eval:rag:offline` 26 suites / 623 tests, `eval:rag:adversarial:offline` 25/25 (3 divergences still pinned), rag.ts 4362/4362; canary pair 32100681177 (`4ea310e48`) -> post-merge dispatch (owner-approved), plus `eval:answer-quality` 30-case before/after and ~10 Gate E questions requested, not run. Note: the 220-word total-length readability ceiling in `scoreAnswerQualityEvalCase` is a known metric confound for A3 (baseline-record §4) | -| S2b | A3: moderate length (if separate review needed) | — | — | Not needed — A3 shipped inside S2 (combined diff stayed reviewable) | — | -| S3 | A4: follow-up suggestion refinement | `claude/rag-a4-follow-ups-` | — | Blocked on S2 merge + its green canary pair | consumes `buildRelatedInformationMenu` from `src/lib/rag/answer-composition.ts` | -| S4 | B0: adversarial fixtures + baseline + register | `claude/packet-s4-adversarial-fixtures-5ho5tp` | #2036 | Merged 2026-08-17 (squash `f5b093291`) | Offline only: `check:rag:adversarial-fixtures` 24 cases / 8 categories / 6 canaries; `eval:rag:offline` 24 suites, 597 tests. Baseline `scripts/fixtures/rag-adversarial-baseline.v1.json` marks the three provider-backed gates `pending_owner_run` | -| S5 | B1+B2: telemetry assessment + offline harness | `claude/s5-rag-telemetry-harness-2wvis7` | #2056 | Merged 2026-08-17 (merge `093f9340c`); post-merge canary run 32049952885 | Offline only: `eval:rag:adversarial:offline` 25/25 (24 cases + canary-free report; 3 divergences pinned in `KNOWN_DIVERGENCES`); B1 gap = `verification_latency_ms` behind `RAG_TELEMETRY_EXTENDED` (default false); canary-absence tests green; 44/44 denominator reconciled | -| S6 | B3: Docling lab benchmark | `claude/packet-s6-docling-lab-d6foa6` | #2057 | Merged 2026-08-17 (merge `5a6418636`) | Offline only: `check:docling-lab` 36 fixtures / 10 hostile / 6 canaries + Gate B template valid; `verify:pr-local` heavy plan failed:(none); contract test 20/20; legacy smoke 46 docs, 10/10 hostile contained, canary-clean report. Verdict is a separate owner dispatch of `docling-lab.yml` | -| S7+ | B4 shadow / B5 Ragas / B6 reranker / B7 DSPy | — | — | Gated — owner decision | — | -| #212 T1–T3 | Runtime row contracts (rag.ts, rag-candidate-sources.ts, src/app/api) — sibling stream sharing `src/lib/rag/**` | — | #1946 / #1981 / #2023 | Merged (T3 squash `440a34f71` 2026-08-17) | see the #212 ledger row; RAG surface complete for the cast class | -| #212 T4 | Runtime row contracts: `worker/main.ts` (11 casts) — sibling stream | `claude/ledger-212-tranche-4-worker-q3y6i4` | #2037 | Merged 2026-08-17 (squash `1726537b7`); #212 closed by reconcile PR #2045 | Governance Preflight complete; audit: 1 inbound cast (claim rows, per-row fail-soft) + 2 read-back param casts contracted, 9 outbound/interop left; closes #212 (inbox `done` queued in the PR) | +| Packet | Scope | Branch | PR | State | Canary / evidence refs | +| ---------- | ----------------------------------------------------------------------------------------------------------------------------------------------------- | ------------------------------------------------ | --------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | +| Guide | Programme guide | `claude/rag-plan-review-guide-vhrls9` | #1895 | Merged 2026-08-13 | docs-only | +| Handover | Multi-session handover + coordination | `claude/rag-plan-review-guide-vhrls9` | #1908 / #2024 | Merged 2026-08-13; coordination layer PR #2024 | docs-only | +| S0 | A1 phase 1: structured fallback diagnostics | `claude/lithium-generation-quality-debug-ji1vce` | #1899 | Merged 2026-08-13 | offline 93/93 focused | +| S1 | A1 phase 2: rung-1 verification-faithfulness fixes | `claude/s1-rag-mitigation-231-86c182` | #2022 | Merged 2026-08-17 (squash `2bd146eed`, landed by content) | 8 pre-fix + 5 post-fix live probes 2026-08-17; offline 583/583; canary pair run 31964560921 (baseline `8f8d111ab`) -> run 32025082010 (`2bd146eed`): recall 1.0/1.0, zero per-case rr regressions, answer gate 44/44 (denominator reconciled by S5; see baseline-record §3); rung-2 measurement in `docs/audit/live-drift-forensics-2026-08.md` §5 | +| S1b | A1 rung 3 (R1): pre-deadline strong routing for dosing class | `claude/s1b-rag-dosing-routing-6u1mik` | #2035 | Merged 2026-08-17 (PR #2035, merge `92f7618`) | canary pair pending: baseline run 32025082010 (`2bd146eed`) -> post-merge dispatch (owner-approved); offline 586/586 + verify:pr-local heavy scope green | +| S1c | A1 residuals R2 + R3: claim-support strictness | `claude/s1c-residuals-r2-r3-4pb1at` | #2052 | Merged 2026-08-17 (merge `b8e774bcd`; follow-up #2063 kept; follow-up #2065 reverted by PR #2088 after canary regression) | canary pair: baseline run 32049952885 -> post run 32052479537 (`084f63799`): recall 1.0/1.0, zero per-case rr regressions, answer gate 44/44 | +| S1d | A1 final-gate gap recovery: hedged cited low-confidence fast answers must recover extractively, not collapse to a citation-free `provider_source_gap` | `claude/s1d-final-gate-gap-recovery-dxgrn2` | #2054 | Merged 2026-08-17 (merge `0bbd64fbc`); landed by content; canary pair green after the #2065 revert (PR #2088) | canary pair 32052479537 -> 32100681177 (`4ea310e48`) green: recall 1.0/1.0, zero per-case rr regressions, answer gate 44/44. Interim post run 32097916649 (`9904fbda8`) was RED on `agitation-im-po-route-short-terms` — bisected live to PR #2065 (S1c follow-up condition-first regex), not S1d; reverted by PR #2088; the confirmation run is 32100681177 | +| G1 | Governance: provenance tag for document-summary rows (Option B) | `claude/g1-rag-document-context-qn9ubx` | #2053 | Merged 2026-08-17 (merge `125e98526`); rows #J912J9 / #0MSNT8 closed at reconcile | no canary (no behaviour change); `document_context` tag at `types.ts` + `rag-row-contracts.ts`, deriveConfidence pinned | +| S2 | A2 + A3: composition menu + moderate length | `claude/s2-rag-composition-7330b0` | #2097 | Merged 2026-08-18 (squash dda4956ff), landed by content; A2 and A3 together; prompt clinical-rag-answer-v19, schema v4, answerSections.maxItems 6 | offline: composition 9/9 + prompt pins 5/5, eval:rag:offline 26 suites / 623 tests, eval:rag:adversarial:offline 25/25 (3 divergences still pinned), rag.ts 4362/4362; canary pair 32100681177 (4ea310e48) -> 32111839806 GREEN (document/content recall 1.0/1.0, zero per-case rr regressions, answer gate clean; one non-blocking 20 s latency advisory on neuroleptic-side-effect-escalation); eval:answer-quality v18 (4ea310e48) vs v19 run 2026-08-18 neutral within nondeterminism (relevance 0.633->0.567 on two nondeterministic timeout/gap cases, targeting 0.409->0.429, readability at ceiling), owner blinded read pending. Note: the 220-word total-length readability ceiling in scoreAnswerQualityEvalCase is a known metric confound for A3 (baseline-record §4) | +| S2b | A3: moderate length (if separate review needed) | — | — | Not needed — A3 shipped inside S2 (combined diff stayed reviewable) | — | +| S3 | A4: follow-up suggestion refinement | `claude/s3-follow-up-suggestions-95e160` | #2108 | PR open 2026-08-18 (menu-derived, evidence-gated, already-answered-suppressed chips in src/lib/answer-follow-up.ts; no ClinicalDashboard change, no new module/field/render block); owner merges | offline only: answer-follow-up 27/27 (menu alignment across all 48 class×intent cells, evidence gate positive + discriminating negatives, suppression), answer-follow-up-chips DOM 6/6 (desktop + phone composer surfaces), answer-composition 9/9 unchanged; no canary — deterministic composition only, generation prompt untouched | +| S4 | B0: adversarial fixtures + baseline + register | `claude/packet-s4-adversarial-fixtures-5ho5tp` | #2036 | Merged 2026-08-17 (squash `f5b093291`) | Offline only: `check:rag:adversarial-fixtures` 24 cases / 8 categories / 6 canaries; `eval:rag:offline` 24 suites, 597 tests. Baseline `scripts/fixtures/rag-adversarial-baseline.v1.json` marks the three provider-backed gates `pending_owner_run` | +| S5 | B1+B2: telemetry assessment + offline harness | `claude/s5-rag-telemetry-harness-2wvis7` | #2056 | Merged 2026-08-17 (merge `093f9340c`); post-merge canary run 32049952885 | Offline only: `eval:rag:adversarial:offline` 25/25 (24 cases + canary-free report; 3 divergences pinned in `KNOWN_DIVERGENCES`); B1 gap = `verification_latency_ms` behind `RAG_TELEMETRY_EXTENDED` (default false); canary-absence tests green; 44/44 denominator reconciled | +| S6 | B3: Docling lab benchmark | `claude/packet-s6-docling-lab-d6foa6` | #2057 | Merged 2026-08-17 (merge `5a6418636`) | Offline only: `check:docling-lab` 36 fixtures / 10 hostile / 6 canaries + Gate B template valid; `verify:pr-local` heavy plan failed:(none); contract test 20/20; legacy smoke 46 docs, 10/10 hostile contained, canary-clean report. Verdict is a separate owner dispatch of `docling-lab.yml` | +| S7+ | B4 shadow / B5 Ragas / B6 reranker / B7 DSPy | — | — | Gated — owner decision | — | +| #212 T1–T3 | Runtime row contracts (rag.ts, rag-candidate-sources.ts, src/app/api) — sibling stream sharing `src/lib/rag/**` | — | #1946 / #1981 / #2023 | Merged (T3 squash `440a34f71` 2026-08-17) | see the #212 ledger row; RAG surface complete for the cast class | +| #212 T4 | Runtime row contracts: `worker/main.ts` (11 casts) — sibling stream | `claude/ledger-212-tranche-4-worker-q3y6i4` | #2037 | Merged 2026-08-17 (squash `1726537b7`); #212 closed by reconcile PR #2045 | Governance Preflight complete; audit: 1 inbound cast (claim rows, per-row fail-soft) + 2 read-back param casts contracted, 9 outbound/interop left; closes #212 (inbox `done` queued in the PR) | Update rule: the session that opens a packet's PR edits its row (branch, PR number, state) in the same PR. A later session updating another packet may also correct stale rows diff --git a/src/lib/answer-follow-up.ts b/src/lib/answer-follow-up.ts index f5a983f02b..1d74ebca41 100644 --- a/src/lib/answer-follow-up.ts +++ b/src/lib/answer-follow-up.ts @@ -1,4 +1,5 @@ -import type { RagAnswer } from "@/lib/types"; +import { buildRelatedInformationMenu, type RelatedInformationMenuKey } from "@/lib/rag/answer-composition"; +import type { AnswerSectionKind, RagAnswer } from "@/lib/types"; // Client-side follow-up context for the single-query answer API. // The /api/answer/stream schema accepts one query string (max 2000 chars), so @@ -64,17 +65,288 @@ export function buildAnswerFollowUpQuery(priorQuery: string | undefined, followU return `Follow-up to "${trimmedPrior.slice(0, budget)}": ${trimmedFollowUp}`; } +// --------------------------------------------------------------------------- +// Follow-up suggestion chips (RAG improvement programme, packet S3 / README §A4) +// +// The chips are deterministic and provider-free, and they are clinical output: +// a chip names a question this corpus is expected to be able to answer. Three +// rules make that claim honest, and each is pinned in tests/answer-follow-up.test.ts: +// +// 1. Candidates come from the S2 related-information menu +// (`buildRelatedInformationMenu`, src/lib/rag/answer-composition.ts) for the +// answer's own (queryClass, intent) cell — the same menu the generation +// prompt was given. This module READS that menu and never changes it. +// 2. Evidence gate: a candidate is emitted only when the retrieved evidence the +// client actually received mentions its subject matter. The client sees a +// bounded ≤900-char snippet per source (trimSourceForClient in +// answer-client-payload.ts), so the gate can under-suggest but never +// over-suggest — the safe direction for a clinical surface. +// 3. Already-answered suppression: a menu kind that the answer already emitted +// as an `answerSections` entry, or whose concept words are already in the +// answer text, is dropped rather than offered back to the reader. +// +// Query classes whose menu is "none" (document_lookup, unsupported_or_general) +// keep the narrow legacy templates — §A2's "stays narrow" rule — but run through +// the same gate and suppression. When nothing survives, the function returns [] +// and no chip row renders; that is the intended conservative outcome, not a bug. +// --------------------------------------------------------------------------- + const maxFollowUpSuggestions = 4; +/** Bounded evidence scan: chips must cost nothing measurable inside the render memo. */ +const maxEvidenceSources = 12; +const maxEvidenceFieldChars = 1200; +const maxEvidenceQuotes = 8; + function normalizeSuggestionKey(value: string) { return value.trim().toLowerCase(); } -function topicLabel(priorQuery: string, answer: RagAnswer) { - const medications = answer.queryAnalysis?.medications ?? []; - const medication = medications.find((item) => item.trim()); - if (medication) return medication.trim(); +/** + * One deterministic follow-up candidate. + * + * `evidenceTerms` are lowercase substrings; at least one must appear in the + * retrieved evidence for the candidate to be offered. `answeredTerms` are the + * narrow concept words that mark the question as already answered by the answer + * body — deliberately a subset of `evidenceTerms`, so a passing mention in a + * source does not suppress a question the answer never addressed. + */ +type FollowUpTemplate = { + kind: AnswerSectionKind; + question: (topic: string) => string; + evidenceTerms: readonly string[]; + answeredTerms: readonly string[]; +}; + +/** + * Follow-up phrasing per menu key, index-aligned with that menu's `items` in + * src/lib/rag/answer-composition.ts. The alignment (same kind at the same index) + * is asserted over all 48 class×intent cells in tests/answer-follow-up.test.ts, so + * a future menu edit fails loudly instead of silently dropping a chip. + */ +const menuFollowUpTemplates: Record, readonly FollowUpTemplate[]> = { + dosing: [ + { + kind: "monitoring_timing", + question: (topic) => `What monitoring is required for ${topic}?`, + evidenceTerms: ["monitor", "bloods", "fbc", "ecg", "egfr", "tft", "serum", "level"], + answeredTerms: ["monitor"], + }, + { + kind: "contraindications_cautions", + question: (topic) => `What cautions or contraindications apply to ${topic}?`, + evidenceTerms: ["contraindicat", "caution", "avoid", "not recommended", "interaction"], + answeredTerms: ["contraindicat", "caution", "avoid"], + }, + { + kind: "escalation_risk", + question: (topic) => `What should trigger stopping or escalating ${topic}?`, + evidenceTerms: ["stop", "withhold", "cease", "discontinu", "escalat", "toxicity"], + answeredTerms: ["stop", "withhold", "cease", "discontinu"], + }, + { + kind: "medication_dose", + question: (topic) => `How is ${topic} dosed in renal or hepatic impairment?`, + evidenceTerms: ["renal", "hepatic", "egfr", "creatinine", "liver", "impairment"], + answeredTerms: ["renal", "hepatic", "egfr", "creatinine"], + }, + ], + escalation: [ + { + kind: "required_actions", + question: (topic) => `What immediate actions are required for ${topic}?`, + evidenceTerms: ["immediate", "urgent", "stop", "withhold", "seek", "first-line"], + answeredTerms: ["immediate", "urgent"], + }, + { + kind: "thresholds", + question: (topic) => `What thresholds trigger those actions for ${topic}?`, + evidenceTerms: ["threshold", "mmol", "exceeds", "above", "below", "score"], + answeredTerms: ["threshold", "mmol/l", "cut-off"], + }, + { + kind: "escalation_risk", + question: (topic) => `Who should be contacted or referred for ${topic}?`, + evidenceTerms: ["refer", "contact", "on-call", "consultant", "specialist", "emergency department", "notify"], + answeredTerms: ["refer", "contact", "on-call", "notify"], + }, + { + kind: "documentation", + question: (topic) => `What should I document for ${topic}?`, + evidenceTerms: ["document", "record", "chart", "form"], + answeredTerms: ["document", "record"], + }, + ], + threshold: [ + { + kind: "thresholds", + question: (topic) => `What are the adjacent thresholds or bands for ${topic}?`, + evidenceTerms: ["threshold", "band", "range", "cut-off", "score", "mmol"], + answeredTerms: ["threshold", "band", "range"], + }, + { + kind: "required_actions", + question: (topic) => `What action is required at each band for ${topic}?`, + evidenceTerms: ["action", "repeat", "manage", "treat", "review"], + answeredTerms: ["action required", "repeat"], + }, + { + kind: "escalation_risk", + question: (topic) => `What should trigger escalation for ${topic}?`, + evidenceTerms: ["escalat", "urgent", "refer", "senior", "emergency"], + answeredTerms: ["escalat"], + }, + ], + comparison: [ + { + kind: "comparison", + question: (topic) => `Which factors decide between the options for ${topic}?`, + evidenceTerms: ["factor", "prefer", "first-line", "choice", "consider"], + answeredTerms: ["factor", "prefer"], + }, + { + kind: "comparison", + question: (topic) => `Where do the sources differ on ${topic}?`, + evidenceTerms: ["differ", "versus", "compared", "whereas", "alternative"], + answeredTerms: ["differ", "conflict"], + }, + { + kind: "required_actions", + question: (topic) => `What switching or washout steps apply to ${topic}?`, + evidenceTerms: ["switch", "washout", "cross-taper", "taper", "swap"], + answeredTerms: ["switch", "washout", "taper"], + }, + ], + management: [ + { + kind: "required_actions", + question: (topic) => `What is the first-line management for ${topic}?`, + evidenceTerms: ["first-line", "first line", "offer", "management", "treat", "initial"], + answeredTerms: ["first-line", "first line"], + }, + { + kind: "documentation", + question: (topic) => `What should I document for ${topic}?`, + evidenceTerms: ["document", "record", "chart", "form"], + answeredTerms: ["document", "record"], + }, + { + kind: "source_gap", + question: (topic) => `What does the indexed guidance not cover for ${topic}?`, + // Only offered when the answer itself reported a gap or conflict; the + // haystack check below is satisfied by that report, not by source prose. + evidenceTerms: ["reported_gap"], + answeredTerms: [], + }, + ], +}; + +/** + * The kinds this module answers for each S2 menu key, in menu order. Exported so + * tests/answer-follow-up.test.ts can assert index-alignment with + * `buildRelatedInformationMenu` across every class×intent cell without reaching + * into the template bodies. + */ +export const followUpTemplateKindsByMenuKey: Readonly< + Record, readonly AnswerSectionKind[]> +> = Object.freeze({ + dosing: menuFollowUpTemplates.dosing.map((template) => template.kind), + escalation: menuFollowUpTemplates.escalation.map((template) => template.kind), + threshold: menuFollowUpTemplates.threshold.map((template) => template.kind), + comparison: menuFollowUpTemplates.comparison.map((template) => template.kind), + management: menuFollowUpTemplates.management.map((template) => template.kind), +}); + +/** + * Candidates for the classes whose S2 menu is deliberately "none". These are the + * pre-S3 narrow templates, kept verbatim in phrasing and now evidence-gated. + */ +const documentLookupTemplates: readonly FollowUpTemplate[] = [ + { + kind: "required_actions", + question: () => "What are the key action points?", + evidenceTerms: ["action", "step", "must", "should", "require"], + answeredTerms: ["key action"], + }, + { + kind: "monitoring_timing", + question: () => "What monitoring or follow-up is documented?", + evidenceTerms: ["monitor", "follow-up", "review"], + answeredTerms: ["monitor", "follow-up"], + }, + { + kind: "contraindications_cautions", + question: () => "Are there any contraindications noted?", + evidenceTerms: ["contraindicat", "caution", "avoid", "not recommended"], + answeredTerms: ["contraindicat"], + }, + { + kind: "required_actions", + question: (topic) => `Summarise the practical steps for ${topic}.`, + evidenceTerms: ["step", "procedure", "process", "plan", "must", "should"], + answeredTerms: [], + }, +]; +const generalTemplates: readonly FollowUpTemplate[] = [ + { + kind: "monitoring_timing", + question: (topic) => `What monitoring is required for ${topic}?`, + evidenceTerms: ["monitor", "review", "level"], + answeredTerms: ["monitor"], + }, + { + kind: "contraindications_cautions", + question: (topic) => `What are the main cautions for ${topic}?`, + evidenceTerms: ["caution", "contraindicat", "avoid", "risk", "adverse"], + answeredTerms: ["caution", "contraindicat"], + }, + { + kind: "documentation", + question: (topic) => `What should I document for ${topic}?`, + evidenceTerms: ["document", "record", "form"], + answeredTerms: ["document", "record"], + }, + { + kind: "required_actions", + question: () => "What would change the management plan?", + evidenceTerms: ["manage", "plan", "treat", "review"], + answeredTerms: [], + }, +]; + +function bounded(value: string | null | undefined): string { + if (!value) return ""; + return value.length > maxEvidenceFieldChars ? value.slice(0, maxEvidenceFieldChars) : value; +} + +/** + * The retrieved evidence this client actually received, lowercased for substring + * matching. Sources carry the bounded client snippet; quote cards and the + * server-derived safety findings add the passages the UI already shows. Query + * analysis is deliberately excluded — canonical terms are query-derived, so + * counting them here would let a query term vouch for corpus coverage it does + * not have. They vouch for the SUBJECT only (see `buildSubjectHaystack`). + */ +function buildEvidenceHaystack(answer: RagAnswer): string { + const parts: string[] = []; + for (const source of (answer.sources ?? []).slice(0, maxEvidenceSources)) { + parts.push(source.title ?? "", source.section_heading ?? "", bounded(source.retrieval_synopsis ?? source.content)); + } + const quotes = [...(answer.quoteCards ?? []), ...(answer.smartPanel?.quotes ?? [])].slice(0, maxEvidenceQuotes); + for (const quote of quotes) parts.push(bounded(quote.quote)); + for (const warning of answer.safetyWarnings ?? []) parts.push(warning.kind, warning.label, bounded(warning.text)); + return parts.join(" \n").toLowerCase(); +} + +/** Where a suggested subject may legitimately come from: evidence, or the resolved query analysis. */ +function buildSubjectHaystack(answer: RagAnswer, evidenceHaystack: string): string { + const analysis = answer.queryAnalysis; + const terms = [...(analysis?.medications ?? []), ...(analysis?.canonicalTerms ?? [])]; + return `${evidenceHaystack} \n${terms.join(" \n").toLowerCase()}`; +} + +function topicLabel(priorQuery: string, answer: RagAnswer) { const canonical = answer.queryAnalysis?.canonicalTerms?.filter((term) => term.trim()) ?? []; if (canonical.length > 0) { const label = canonical.slice(0, 3).join(" "); @@ -102,82 +374,60 @@ function isShortContinuationQuery(query: string) { } /** - * Pick the clinical topic embedded in follow-up suggestion chips. Short - * continuation questions ("what about renal impairment?") should anchor on the - * thread's opening question, not echo the latest follow-up phrasing. + * The thread question a chip's topic should be read from. Short continuation + * questions ("what about renal impairment?") anchor on the thread's opening + * question, not the latest follow-up phrasing. */ -function resolveSuggestionTopicAnchor(latestQuery: string, priorQueries: string[], answer: RagAnswer) { - const medications = answer.queryAnalysis?.medications ?? []; - const medication = medications.find((item) => item.trim()); - if (medication) return medication.trim(); - - const threadQueries = priorQueries.map((query) => query.trim()).filter(Boolean); - const firstQuery = threadQueries[0]; - if (firstQuery && firstQuery !== latestQuery.trim() && isShortContinuationQuery(latestQuery)) { - return firstQuery; - } - return latestQuery.trim(); -} - -function medicationFollowUpTemplates(topic: string) { - return [ - "What about renal impairment?", - "What monitoring is required?", - "What are the elderly dosing considerations?", - "What about pregnancy or breastfeeding?", - "Are there important drug interactions?", - `What cautions apply to ${topic}?`, - ]; +function resolveAnchorQuery(latestQuery: string, priorQueries: string[]) { + const trimmedLatest = latestQuery.trim(); + const firstQuery = priorQueries.map((query) => query.trim()).find(Boolean); + if (firstQuery && firstQuery !== trimmedLatest && isShortContinuationQuery(trimmedLatest)) return firstQuery; + return trimmedLatest; } -function thresholdFollowUpTemplates(topic: string) { - return [ - "What should trigger escalation?", - "What are the alternative thresholds?", - `When should I repeat ${topic}?`, - "What monitoring supports this threshold?", - ]; -} - -function comparisonFollowUpTemplates(topic: string) { - return [ - "Which option is preferred in pregnancy?", - "Which option needs less monitoring?", - `What are the key differences for ${topic}?`, - "When would you choose the alternative?", - ]; +/** + * Prefer a subject the clinician actually named. Medications on the analysis can + * be answer-derived rather than asked-for — an agitation question whose answer + * lists olanzapine must not produce "…for olanzapine?" chips — so a term that + * appears in the anchor question wins, and a bare medication is only the + * fallback. + */ +function subjectFromAnalysis(anchorQuery: string, answer: RagAnswer): string | null { + const normalizedAnchor = anchorQuery.toLowerCase(); + const medications = (answer.queryAnalysis?.medications ?? []).map((term) => term.trim()).filter(Boolean); + const canonical = (answer.queryAnalysis?.canonicalTerms ?? []).map((term) => term.trim()).filter(Boolean); + for (const term of [...medications, ...canonical]) { + const at = normalizedAnchor.indexOf(term.toLowerCase()); + if (at < 0) continue; + // Read the subject back out of the question so the chip keeps the + // clinician's own casing ("ADHD", not the normalised "adhd"). + return anchorQuery.slice(at, at + term.length); + } + return medications[0] ?? null; } -function genericFollowUpTemplates(topic: string) { - return [ - `What monitoring is required for ${topic}?`, - `What are the main cautions for ${topic}?`, - `What should I document for ${topic}?`, - "What would change the management plan?", - ]; +/** + * Which composition menu applies. Class-carrying answers use the S2 menu + * verbatim; the query-shape fallbacks only cover answers that reached the client + * without a query analysis (cached or degraded payloads). + */ +function resolveMenuKey(anchorQuery: string, answer: RagAnswer): RelatedInformationMenuKey { + const queryClass = answer.queryClass ?? answer.queryAnalysis?.queryClass; + if (queryClass) { + return buildRelatedInformationMenu(queryClass, answer.queryAnalysis?.intent ?? "general").key; + } + if (answer.queryAnalysis?.comparisonIntent) return "comparison"; + if (answer.queryAnalysis?.documentTitleIntent) return "none"; + if (/\b(dose|dosing|mg|monitor|medication|drug)\b/i.test(anchorQuery)) return "dosing"; + if (/\b(threshold|level|cut[- ]?off|range)\b/i.test(anchorQuery)) return "threshold"; + return "none"; } -function templatesForAnswer(priorQuery: string, answer: RagAnswer) { - const topic = topicLabel(priorQuery, answer); +function templatesForMenuKey(menuKey: RelatedInformationMenuKey, answer: RagAnswer): readonly FollowUpTemplate[] { + if (menuKey !== "none") return menuFollowUpTemplates[menuKey]; const queryClass = answer.queryClass ?? answer.queryAnalysis?.queryClass; - if (queryClass === "medication_dose_risk" || /\b(dose|dosing|mg|monitor|medication|drug)\b/i.test(priorQuery)) { - return medicationFollowUpTemplates(topic); - } - if (queryClass === "table_threshold" || /\b(threshold|level|cut[- ]?off|range)\b/i.test(priorQuery)) { - return thresholdFollowUpTemplates(topic); - } - if (queryClass === "comparison" || answer.queryAnalysis?.comparisonIntent) { - return comparisonFollowUpTemplates(topic); - } - if (queryClass === "document_lookup" || answer.queryAnalysis?.documentTitleIntent) { - return [ - "What are the key action points?", - "What monitoring or follow-up is documented?", - "Are there any contraindications noted?", - `Summarise the practical steps for ${topic}.`, - ]; - } - return genericFollowUpTemplates(topic); + if (queryClass === "document_lookup" || answer.queryAnalysis?.documentTitleIntent) return documentLookupTemplates; + return generalTemplates; } function gapFollowUpTemplates(answer: RagAnswer) { @@ -194,7 +444,12 @@ function gapFollowUpTemplates(answer: RagAnswer) { /** * Build short follow-up question chips for the latest answer turn. - * Suggestions avoid repeating questions already asked in the thread. + * + * Candidates come from the S2 composition menu for this answer's class/intent, + * must be supported by the retrieved evidence, and are dropped when the answer + * already covers them. Returns [] when nothing survives — the chip row then does + * not render, which is preferable to naming a question this corpus cannot answer. + * Suggestions also avoid repeating questions already asked in the thread. */ export function buildAnswerFollowUpSuggestions( priorQuery: string, @@ -204,18 +459,54 @@ export function buildAnswerFollowUpSuggestions( const trimmedPrior = priorQuery.trim(); if (!trimmedPrior) return []; - const topicQuery = resolveSuggestionTopicAnchor(trimmedPrior, priorQueries, answer); + const anchorQuery = resolveAnchorQuery(trimmedPrior, priorQueries); + const evidenceHaystack = buildEvidenceHaystack(answer); + const subjectHaystack = buildSubjectHaystack(answer, evidenceHaystack); + + const topic = subjectFromAnalysis(anchorQuery, answer) ?? topicLabel(anchorQuery, answer); + const topicSupported = significantTokens(topic).some((token) => subjectHaystack.includes(token)); + const seen = new Set(priorQueries.map(normalizeSuggestionKey)); seen.add(normalizeSuggestionKey(trimmedPrior)); const suggestions: string[] = []; - for (const candidate of [...gapFollowUpTemplates(answer), ...templatesForAnswer(topicQuery, answer)]) { + const push = (candidate: string) => { const normalized = normalizeSuggestionKey(candidate); - if (!normalized || seen.has(normalized)) continue; - if (suggestions.some((item) => normalizeSuggestionKey(item) === normalized)) continue; + if (!normalized || seen.has(normalized)) return; suggestions.push(candidate); seen.add(normalized); + }; + + // Reported gaps and conflicts are evidence by construction: the answer itself + // said the sources disagree or fall short. + for (const gap of gapFollowUpTemplates(answer)) { if (suggestions.length >= maxFollowUpSuggestions) break; + push(gap); } + const hasReportedGap = (answer.conflictsOrGaps ?? answer.smartPanel?.conflictsOrGaps ?? []).length > 0; + const answerText = (answer.answer ?? "").toLowerCase(); + const emittedSectionKinds = new Set( + (answer.answerSections ?? []).map((section) => section.kind).filter((kind): kind is AnswerSectionKind => !!kind), + ); + + if (!topicSupported) return suggestions; + + const menuKey = resolveMenuKey(anchorQuery, answer); + for (const template of templatesForMenuKey(menuKey, answer)) { + if (suggestions.length >= maxFollowUpSuggestions) break; + // (2) evidence gate. + const supported = template.evidenceTerms.some((term) => + term === "reported_gap" ? hasReportedGap : evidenceHaystack.includes(term), + ); + if (!supported) continue; + // (3) already-answered suppression: an emitted section of this kind, the + // answer body's own words, or — for the source-gap item — a gap chip that is + // already asking the specific question. + if (emittedSectionKinds.has(template.kind)) continue; + if (template.kind === "source_gap" && suggestions.length > 0) continue; + if (template.answeredTerms.some((term) => answerText.includes(term))) continue; + push(template.question(topic)); + } + return suggestions; } diff --git a/tests/answer-follow-up-chips.dom.test.tsx b/tests/answer-follow-up-chips.dom.test.tsx new file mode 100644 index 0000000000..d906a44091 --- /dev/null +++ b/tests/answer-follow-up-chips.dom.test.tsx @@ -0,0 +1,162 @@ +import { readFileSync } from "node:fs"; +import path from "node:path"; + +import { render, screen, within } from "@testing-library/react"; +import userEvent from "@testing-library/user-event"; +import { describe, expect, it, vi } from "vitest"; + +import { AnswerFollowUpSuggestions } from "@/components/clinical-dashboard/answer-follow-up-suggestions"; +import { buildAnswerFollowUpSuggestions } from "@/lib/answer-follow-up"; +import type { RagAnswer, SearchResult } from "@/lib/types"; + +// Packet S3 / README §A4. The follow-up chips render on TWO surfaces — the +// desktop answer surface (`answer-result-surface.tsx`, `hidden sm:block`) and the +// phone composer dock (`master-search-header.tsx`, `sm:hidden`) — and both are fed +// by the same `buildAnswerFollowUpSuggestions` output. This file renders the chip +// component with each call site's exact props so a menu/evidence change is proved +// end-to-end in the DOM, and pins the two call sites so the wiring cannot drift +// away from that function (the same source-contract pattern as +// tests/phone-dock-addon-contract.test.ts). + +function read(relativePath: string): string { + return readFileSync(path.resolve(process.cwd(), relativePath), "utf8"); +} + +const desktopSurface = read("src/components/clinical-dashboard/answer-result-surface.tsx"); +const phoneHeader = read("src/components/clinical-dashboard/master-search-header.tsx"); +const dashboard = read("src/components/ClinicalDashboard.tsx"); + +const source: SearchResult = { + id: "chunk-0", + document_id: "doc-1", + title: "WA Lithium Prescribing Guideline", + file_name: "lithium.pdf", + page_number: 4, + chunk_index: 0, + section_heading: "Prescribing", + content: + "Monitor serum lithium levels after each dose change. Cautions and contraindications are listed below. " + + "Reduce the dose in renal impairment.", + image_ids: [], + images: [], + similarity: 0.71, +}; + +const answer: RagAnswer = { + answer: "Start lithium 400 mg nocte and titrate against response.", + grounded: true, + confidence: "high", + citations: [], + sources: [source], + queryClass: "medication_dose_risk", + queryAnalysis: { + originalQuery: "lithium dosing", + normalizedQuery: "lithium dosing", + queryClass: "medication_dose_risk", + intent: "drug_dosing", + confidence: 0.9, + reasons: [], + canonicalTerms: ["lithium"], + expandedTerms: [], + typoCorrections: [], + medications: ["lithium"], + acronyms: [], + thresholdTerms: [], + documentTitleTerms: [], + queryRewrite: { normalizedQuery: "lithium dosing", searchQuery: "lithium dosing", expansions: [], reasons: [] }, + documentTitleIntent: false, + comparisonIntent: false, + freshnessNeed: false, + needsVisualEvidence: false, + needsSynthesis: false, + needsClassifierFallback: false, + }, +}; + +const suggestions = buildAnswerFollowUpSuggestions("lithium dosing", answer, ["lithium dosing"]); + +/** The desktop call site: default test id, wrap layout (answer-result-surface.tsx). */ +function renderDesktop(onPick: (suggestion: string) => void, disabled = false) { + return render(); +} + +/** The phone call site: composer test id, scroll layout (master-search-header.tsx). */ +function renderPhone(onPick: (suggestion: string) => void, disabled = false) { + return render( + , + ); +} + +describe("answer follow-up chips · menu-derived output reaches both surfaces", () => { + it("renders the evidence-supported menu chips as wired buttons on the desktop answer surface", async () => { + const onPick = vi.fn(); + renderDesktop(onPick); + + const row = screen.getByTestId("answer-follow-up-suggestions"); + const chips = within(row).getAllByRole("button"); + expect(chips.map((chip) => chip.textContent)).toEqual(suggestions); + expect(suggestions).toContain("What monitoring is required for lithium?"); + + await userEvent.click(chips[0]); + expect(onPick).toHaveBeenCalledWith(suggestions[0]); + }); + + it("renders the same chips in the phone composer dock row", async () => { + const onPick = vi.fn(); + renderPhone(onPick); + + const row = screen.getByTestId("answer-composer-follow-up-suggestions"); + expect(row).toHaveClass("answer-suggestion-row-scroll"); + const chips = within(row).getAllByRole("button"); + expect(chips.map((chip) => chip.textContent)).toEqual(suggestions); + + await userEvent.click(chips[chips.length - 1]); + expect(onPick).toHaveBeenCalledWith(suggestions[suggestions.length - 1]); + }); + + it("never renders a chip whose subject the retrieved evidence did not surface", () => { + renderDesktop(vi.fn()); + const row = screen.getByTestId("answer-follow-up-suggestions"); + + // The source snippet says nothing about stopping, escalation or toxicity, so + // the dosing menu's escalation item must not reach either surface. + expect(within(row).queryByText(/stopping or escalating/i)).toBeNull(); + }); + + it("renders no chip row at all when the gate leaves nothing to suggest", () => { + const { container } = render( + , + ); + expect(container).toBeEmptyDOMElement(); + }); + + it("keeps chips inert while a request is in flight", () => { + renderPhone(vi.fn(), true); + for (const chip of within(screen.getByTestId("answer-composer-follow-up-suggestions")).getAllByRole("button")) { + expect(chip).toBeDisabled(); + } + }); +}); + +describe("answer follow-up chips · call-site wiring contract", () => { + it("keeps both surfaces fed by buildAnswerFollowUpSuggestions", () => { + expect(dashboard).toContain("buildAnswerFollowUpSuggestions"); + // Desktop: the staged answer surface receives the memoised list. + expect(dashboard).toContain("followUpSuggestions={answerFollowUpSuggestions}"); + // Phone: the composer dock receives the same list in answer mode. + expect(dashboard).toContain('composerFollowUpSuggestions={searchMode === "answer" ? answerFollowUpSuggestions'); + expect(desktopSurface).toContain("suggestions={followUpSuggestions}"); + expect(phoneHeader).toContain("suggestions={composerFollowUpSuggestions}"); + expect(phoneHeader).toContain('testId="answer-composer-follow-up-suggestions"'); + }); +}); diff --git a/tests/answer-follow-up.test.ts b/tests/answer-follow-up.test.ts index 39c01bacbe..cd7a3bd671 100644 --- a/tests/answer-follow-up.test.ts +++ b/tests/answer-follow-up.test.ts @@ -1,6 +1,12 @@ import { describe, expect, it } from "vitest"; -import { buildAnswerFollowUpQuery, buildAnswerFollowUpSuggestions } from "@/lib/answer-follow-up"; +import { + buildAnswerFollowUpQuery, + buildAnswerFollowUpSuggestions, + followUpTemplateKindsByMenuKey, +} from "@/lib/answer-follow-up"; +import { buildRelatedInformationMenu } from "@/lib/rag/answer-composition"; +import type { AnswerSection, ClinicalQueryIntent, RagAnswer, RagQueryClass, SearchResult } from "@/lib/types"; describe("buildAnswerFollowUpQuery", () => { it("returns the follow-up unchanged when there is no prior question", () => { @@ -50,99 +56,404 @@ describe("buildAnswerFollowUpQuery", () => { }); }); -describe("buildAnswerFollowUpSuggestions", () => { - const medicationAnswer = { - answer: "Start low and monitor levels.", +// --------------------------------------------------------------------------- +// Follow-up chips (packet S3 / README §A4). The three contracts under test are: +// candidates come from the S2 composition menu, every candidate must be supported +// by retrieved evidence, and anything the answer already covered is suppressed. +// --------------------------------------------------------------------------- + +const allQueryClasses: RagQueryClass[] = [ + "document_lookup", + "table_threshold", + "medication_dose_risk", + "comparison", + "broad_summary", + "unsupported_or_general", +]; + +const allIntents: ClinicalQueryIntent[] = [ + "definition", + "protocol", + "drug_dosing", + "escalation_risk", + "document_lookup", + "comparison", + "broad_summary", + "general", +]; + +function evidenceSource(content: string, index = 0): SearchResult { + return { + id: `chunk-${index}`, + document_id: "doc-1", + title: "WA Lithium Prescribing Guideline", + file_name: "lithium.pdf", + page_number: 4, + chunk_index: index, + section_heading: "Prescribing", + content, + image_ids: [], + images: [], + similarity: 0.7, + }; +} + +/** Mentions every menu kind's subject matter, so a chip that is dropped was dropped for another reason. */ +const broadEvidence = evidenceSource( + "Monitor serum lithium levels and renal function. Reduce the dose in renal or hepatic impairment. " + + "Cautions and contraindications are listed below; avoid in severe impairment. Withhold or stop lithium and " + + "escalate if toxicity is suspected: seek urgent review, refer to the emergency department and contact the " + + "on-call consultant. Thresholds above 1.2 mmol/L require immediate action. Document the decision in the record " + + "and use the shared-care form. Offer first-line management and consider switching or a washout where the " + + "sources differ on the preferred option.", +); + +function analysisFor( + query: string, + queryClass: RagQueryClass, + intent: ClinicalQueryIntent, + overrides: Partial> = {}, +): NonNullable { + return { + originalQuery: query, + normalizedQuery: query, + queryClass, + intent, + confidence: 0.9, + reasons: [], + canonicalTerms: ["lithium"], + expandedTerms: [], + typoCorrections: [], + medications: ["lithium"], + acronyms: [], + thresholdTerms: [], + documentTitleTerms: [], + queryRewrite: { normalizedQuery: query, searchQuery: query, expansions: [], reasons: [] }, + documentTitleIntent: false, + comparisonIntent: false, + freshnessNeed: false, + needsVisualEvidence: false, + needsSynthesis: false, + needsClassifierFallback: false, + ...overrides, + }; +} + +function answerFor( + options: { + query?: string; + queryClass?: RagQueryClass; + intent?: ClinicalQueryIntent; + answer?: string; + sources?: SearchResult[]; + sections?: AnswerSection[]; + analysisOverrides?: Partial>; + } = {}, +): RagAnswer { + const query = options.query ?? "lithium dosing"; + const queryClass = options.queryClass ?? "medication_dose_risk"; + return { + answer: options.answer ?? "Start low and titrate against response.", grounded: true, confidence: "high", citations: [], - sources: [], - queryClass: "medication_dose_risk", - queryAnalysis: { - originalQuery: "lithium dosing", - normalizedQuery: "lithium dosing", - queryClass: "medication_dose_risk", - intent: "drug_dosing", - confidence: 0.9, - reasons: [], - canonicalTerms: ["lithium"], - expandedTerms: [], - typoCorrections: [], - medications: ["lithium"], - acronyms: [], - thresholdTerms: [], - documentTitleTerms: [], - queryRewrite: { - normalizedQuery: "lithium dosing", - searchQuery: "lithium dosing", - expansions: [], - reasons: [], + sources: options.sources ?? [broadEvidence], + queryClass, + queryAnalysis: analysisFor(query, queryClass, options.intent ?? "drug_dosing", options.analysisOverrides), + ...(options.sections ? { answerSections: options.sections } : {}), + }; +} + +describe("buildAnswerFollowUpSuggestions · composition menu", () => { + it("derives dosing chips from the medication_dose_risk / drug_dosing menu, in menu order", () => { + const suggestions = buildAnswerFollowUpSuggestions("lithium dosing", answerFor(), ["lithium dosing"]); + + expect(suggestions).toEqual([ + "What monitoring is required for lithium?", + "What cautions or contraindications apply to lithium?", + "What should trigger stopping or escalating lithium?", + "How is lithium dosed in renal or hepatic impairment?", + ]); + }); + + it("follows the intent refinement: escalation_risk swaps the dosing menu for the escalation menu", () => { + const suggestions = buildAnswerFollowUpSuggestions( + "lithium toxicity action", + answerFor({ query: "lithium toxicity action", intent: "escalation_risk" }), + ["lithium toxicity action"], + ); + + expect(suggestions).toEqual([ + "What immediate actions are required for lithium?", + "What thresholds trigger those actions for lithium?", + "Who should be contacted or referred for lithium?", + "What should I document for lithium?", + ]); + expect(suggestions.some((item) => /renal or hepatic/i.test(item))).toBe(false); + }); + + it("keeps document_lookup narrow — the class whose S2 menu is deliberately none", () => { + const suggestions = buildAnswerFollowUpSuggestions( + "Summarise the discharge policy", + answerFor({ + query: "Summarise the discharge policy", + queryClass: "document_lookup", + intent: "document_lookup", + answer: "The policy sets out planning and follow-up requirements.", + sources: [ + evidenceSource( + "Discharge planning starts at admission. Staff must reconcile medicines and follow the steps below; " + + "contraindications to early discharge are noted.", + ), + ], + analysisOverrides: { canonicalTerms: ["discharge"], medications: [], documentTitleIntent: true }, + }), + ["Summarise the discharge policy"], + ); + + expect(suggestions).toEqual([ + "What are the key action points?", + "Are there any contraindications noted?", + "Summarise the practical steps for discharge.", + ]); + }); + + it("answers every menu kind the S2 menu offers, index-aligned, across all 48 class×intent cells", () => { + for (const queryClass of allQueryClasses) { + for (const intent of allIntents) { + const menu = buildRelatedInformationMenu(queryClass, intent); + if (menu.key === "none") { + expect(menu.items).toHaveLength(0); + continue; + } + const kinds = followUpTemplateKindsByMenuKey[menu.key]; + expect(kinds, `no follow-up templates for menu key ${menu.key}`).toBeDefined(); + expect(kinds).toEqual(menu.items.map((item) => item.kind)); + } + } + }); + + it("falls back to the query shape when a cached payload carries no query analysis", () => { + const suggestions = buildAnswerFollowUpSuggestions( + "lithium dosing", + { + answer: "Start low and titrate against response.", + grounded: true, + confidence: "high", + citations: [], + sources: [broadEvidence], }, - documentTitleIntent: false, - comparisonIntent: false, - freshnessNeed: false, - needsVisualEvidence: false, - needsSynthesis: false, - needsClassifierFallback: false, - }, - } satisfies import("@/lib/types").RagAnswer; - - it("returns medication follow-up suggestions for the latest turn", () => { - const suggestions = buildAnswerFollowUpSuggestions("lithium dosing", medicationAnswer); - expect(suggestions.length).toBeGreaterThan(0); - expect(suggestions.some((item) => /renal impairment/i.test(item))).toBe(true); + ["lithium dosing"], + ); + + expect(suggestions).toContain("What monitoring is required for lithium dosing?"); }); +}); - it("avoids repeating questions already asked in the thread", () => { - const suggestions = buildAnswerFollowUpSuggestions("lithium dosing", medicationAnswer, [ +describe("buildAnswerFollowUpSuggestions · evidence gate", () => { + it("drops a candidate whose subject the retrieved evidence never mentions", () => { + const withoutRenalOrEscalation = buildAnswerFollowUpSuggestions( "lithium dosing", - "What about renal impairment?", + answerFor({ + sources: [ + evidenceSource( + "Lithium is usually given once daily at night. Monitor serum lithium levels after each change. " + + "Cautions are listed in the appendix.", + ), + ], + }), + ["lithium dosing"], + ); + + // Same class, same menu — only the evidence changed. + expect(withoutRenalOrEscalation).toEqual([ + "What monitoring is required for lithium?", + "What cautions or contraindications apply to lithium?", ]); - expect(suggestions.some((item) => /renal impairment/i.test(item))).toBe(false); + expect(withoutRenalOrEscalation.some((item) => /renal or hepatic/i.test(item))).toBe(false); + }); + + it("returns no chips at all when the answer carries no retrieved evidence", () => { + expect(buildAnswerFollowUpSuggestions("lithium dosing", answerFor({ sources: [] }), ["lithium dosing"])).toEqual( + [], + ); + }); + + it("does not let query-derived canonical terms vouch for corpus coverage", () => { + // "monitoring" is in the query analysis but in no source; the monitoring chip + // must not appear on that basis alone. + const suggestions = buildAnswerFollowUpSuggestions( + "lithium monitoring", + answerFor({ + query: "lithium monitoring", + sources: [evidenceSource("Lithium is a mood stabiliser used in bipolar disorder.")], + analysisOverrides: { canonicalTerms: ["lithium", "monitoring"] }, + }), + ["lithium monitoring"], + ); + + expect(suggestions).toEqual([]); + }); + + it("accepts server-derived safety findings and quote cards as evidence", () => { + const suggestions = buildAnswerFollowUpSuggestions( + "lithium dosing", + { + ...answerFor({ sources: [evidenceSource("Lithium is a mood stabiliser used in bipolar disorder.")] }), + safetyWarnings: [ + { + id: "warn-1", + kind: "monitoring", + label: "Monitoring required", + text: "Monitor serum lithium levels every six months.", + citation: { + chunk_id: "chunk-0", + document_id: "doc-1", + title: "WA Lithium Prescribing Guideline", + file_name: "lithium.pdf", + page_number: 4, + chunk_index: 0, + }, + href: "/documents/doc-1", + }, + ], + }, + ["lithium dosing"], + ); + + expect(suggestions).toEqual(["What monitoring is required for lithium?"]); }); - it("anchors suggestion topics on the opening question after a short follow-up turn", () => { - const answerWithoutMedicationHint = { - ...medicationAnswer, - queryAnalysis: { - ...medicationAnswer.queryAnalysis, - medications: [], + it("returns nothing when the subject itself is unsupported by evidence or analysis", () => { + const suggestions = buildAnswerFollowUpSuggestions( + "quetiapine dosing", + { + answer: "No indexed guidance was found.", + grounded: false, + confidence: "low", + citations: [], + sources: [evidenceSource("Monitor renal function and document the review.")], }, - } satisfies import("@/lib/types").RagAnswer; + ["quetiapine dosing"], + ); + + expect(suggestions).toEqual([]); + }); +}); - const suggestions = buildAnswerFollowUpSuggestions("what about renal impairment?", answerWithoutMedicationHint, [ +describe("buildAnswerFollowUpSuggestions · already-answered suppression", () => { + it("skips the follow-up for a kind the answer already emitted as a section", () => { + const suggestions = buildAnswerFollowUpSuggestions( "lithium dosing", - "what about renal impairment?", + answerFor({ + sections: [ + { + heading: "Monitoring and timing", + body: "Check levels five days after each change.", + citation_chunk_ids: ["chunk-0"], + kind: "monitoring_timing", + }, + ], + }), + ["lithium dosing"], + ); + + expect(suggestions.some((item) => /monitoring is required/i.test(item))).toBe(false); + // The rest of the menu still comes through — suppression is per kind. + expect(suggestions).toContain("What cautions or contraindications apply to lithium?"); + }); + + it("skips a follow-up the answer body already covers", () => { + const suggestions = buildAnswerFollowUpSuggestions( + "lithium dosing", + answerFor({ answer: "Monitor serum lithium levels five to seven days after every dose change." }), + ["lithium dosing"], + ); + + expect(suggestions.some((item) => /monitoring is required/i.test(item))).toBe(false); + expect(suggestions[0]).toBe("What cautions or contraindications apply to lithium?"); + }); + + it("keeps a kind whose words appear only in the sources, not in the answer", () => { + const suggestions = buildAnswerFollowUpSuggestions( + "lithium dosing", + answerFor({ answer: "Start low and titrate against response." }), + ["lithium dosing"], + ); + + expect(suggestions).toContain("What monitoring is required for lithium?"); + }); +}); + +describe("buildAnswerFollowUpSuggestions · thread and shape rules", () => { + it("puts reported gaps first and still respects the four-chip cap", () => { + const suggestions = buildAnswerFollowUpSuggestions( + "lithium dosing", + { + ...answerFor(), + conflictsOrGaps: [ + { type: "gap", message: "Paediatric dosing is not covered." }, + { type: "conflict", message: "The two guidelines disagree on the target level." }, + ], + }, + ["lithium dosing"], + ); + + expect(suggestions).toHaveLength(4); + expect(suggestions[0]).toBe("What does the source say about paediatric dosing is not covered?"); + expect(suggestions[1]).toBe("What does the source say about the two guidelines disagree on the target level?"); + expect(suggestions[2]).toBe("What monitoring is required for lithium?"); + }); + + it("avoids repeating questions already asked in the thread", () => { + const suggestions = buildAnswerFollowUpSuggestions("lithium dosing", answerFor(), [ + "lithium dosing", + "What monitoring is required for lithium?", ]); + expect(suggestions.some((item) => /monitoring is required/i.test(item))).toBe(false); + expect(suggestions).toContain("What cautions or contraindications apply to lithium?"); + }); + + it("anchors the subject on the opening question after a short follow-up turn", () => { + const suggestions = buildAnswerFollowUpSuggestions( + "what about renal impairment?", + answerFor({ analysisOverrides: { medications: [] } }), + ["lithium dosing", "what about renal impairment?"], + ); + expect(suggestions.length).toBeGreaterThan(0); expect(suggestions.every((item) => !/for what about renal impairment/i.test(item))).toBe(true); - expect(suggestions.some((item) => /lithium dosing|What monitoring is required\?/i.test(item))).toBe(true); - }); - - it("uses a concise topic label for long first-turn questions", () => { - const clozapineQuestion = "What clozapine monitoring items are shown in the table image?"; - const tableAnswer = { - answer: "The synthetic clozapine table image highlights core monitoring domains.", - grounded: true, - confidence: "high", - citations: [], - sources: [], - queryClass: "document_lookup", - queryAnalysis: { - ...medicationAnswer.queryAnalysis, - originalQuery: clozapineQuestion, - normalizedQuery: clozapineQuestion, - queryClass: "document_lookup", - medications: [], - canonicalTerms: [], - }, - } satisfies import("@/lib/types").RagAnswer; + expect(suggestions).toContain("What monitoring is required for lithium?"); + }); - const suggestions = buildAnswerFollowUpSuggestions(clozapineQuestion, tableAnswer, [clozapineQuestion]); + it("reads the subject back out of the question so acronyms keep their casing", () => { + const suggestions = buildAnswerFollowUpSuggestions( + "ADHD medication monitoring", + answerFor({ + query: "ADHD medication monitoring", + queryClass: "broad_summary", + intent: "protocol", + answer: "The guideline sets out baseline and ongoing checks.", + sources: [ + evidenceSource("Offer first-line treatment and document the review on the shared-care form for ADHD."), + ], + analysisOverrides: { canonicalTerms: ["adhd"], medications: [] }, + }), + ["ADHD medication monitoring"], + ); - expect(suggestions.length).toBeGreaterThan(0); - expect(suggestions.every((item) => !/for What clozapine monitoring items/i.test(item))).toBe(true); - expect(suggestions.some((item) => /clozapine/i.test(item))).toBe(true); + expect(suggestions).toEqual(["What is the first-line management for ADHD?", "What should I document for ADHD?"]); + }); + + it("is deterministic — the same answer always yields the same chips", () => { + const answer = answerFor(); + expect(buildAnswerFollowUpSuggestions("lithium dosing", answer, ["lithium dosing"])).toEqual( + buildAnswerFollowUpSuggestions("lithium dosing", answer, ["lithium dosing"]), + ); + }); + + it("returns nothing without a prior question", () => { + expect(buildAnswerFollowUpSuggestions(" ", answerFor())).toEqual([]); }); });