diff --git a/.claude/hooks/issues-surface.sh b/.claude/hooks/issues-surface.sh index dab08f83d0..9ff772dbb4 100755 --- a/.claude/hooks/issues-surface.sh +++ b/.claude/hooks/issues-surface.sh @@ -1,11 +1,10 @@ #!/usr/bin/env bash # SessionStart hook — surface the outstanding-work memory into context. # -# Reads docs/outstanding-issues.md (the /issues ledger) and prints a compact, -# glanceable summary of the OPEN items so every session starts already aware of -# what is outstanding. When the trigger is a context reset (compact / resume / -# clear) it also emits a reminder to run `/issues capture` — that is the moment -# a session's in-flight follow-ups are most likely to be lost. +# Reads docs/outstanding-issues.md (the universal /issues ledger) and prints the +# ordered recommended tasks plus open-item counts so every session starts with +# the same repository-wide priorities. When the trigger is a context reset +# (compact / resume / clear) it also emits a reminder to run `/issues capture`. # # Contract: READ-ONLY. Never writes, never commits, never fails a session — it # always exits 0, and every step is guarded so a parse error just yields less @@ -43,9 +42,27 @@ rows="$(awk ' } ' "$ledger" 2>/dev/null || true)" +# --- parse the ordered recommended execution queue -------------------------- +# Emit "ORDERIDACUITYWHENESTIMATE" per recommended row. +recommended="$(awk ' + /^## Recommended execution queue/ { inrecommended=1; next } + /^## / { if (inrecommended) inrecommended=0 } + inrecommended && /^\|[[:space:]]*[0-9]+[[:space:]]*\|/ { + n=split($0, c, "|") + order=c[2]; id=c[3]; acuity=c[5]; timing=c[6]; estimate=c[7] + gsub(/^[ \t]+|[ \t]+$/, "", order) + gsub(/^[ \t]+|[ \t]+$/, "", id) + gsub(/^[ \t]+|[ \t]+$/, "", acuity) + gsub(/^[ \t]+|[ \t]+$/, "", timing) + gsub(/^[ \t]+|[ \t]+$/, "", estimate) + printf "%s\t%s\t%s\t%s\t%s\n", order, id, acuity, timing, estimate + } +' "$ledger" 2>/dev/null || true)" + total="$(printf '%s' "$rows" | grep -c . || true)" -if [ "${total:-0}" -eq 0 ]; then - echo "[issues] Outstanding-work memory (docs/outstanding-issues.md): no open items. Record one with /issues add …" +recommended_total="$(printf '%s' "$recommended" | grep -c . || true)" +if [ "${total:-0}" -eq 0 ] && [ "${recommended_total:-0}" -eq 0 ]; then + echo "[issues] Universal task ledger (docs/outstanding-issues.md): no recommended or open items. Record one with /issues add …" exit 0 fi @@ -54,7 +71,25 @@ count() { printf '%s' "$1" | grep -c . || true; } p1="$(group P1)"; p2="$(group P2)"; p3="$(group P3)" c1="$(count "$p1")"; c2="$(count "$p2")"; c3="$(count "$p3")" -echo "[issues] Outstanding-work memory — ${total} open (${c1}×P1, ${c2}×P2, ${c3}×P3). Source of truth: docs/outstanding-issues.md · read the full list back with /issues." +echo "[issues] Universal task ledger — ${recommended_total} recommended · ${total} open (${c1}×P1, ${c2}×P2, ${c3}×P3). Source of truth: docs/outstanding-issues.md · read the full ledger with /issues." + +print_recommended() { # $1=max-to-list + local limit="$1" shown=0 more=0 order id acuity timing estimate + [ -z "$recommended" ] && return 0 + while IFS=$'\t' read -r order id acuity timing estimate; do + [ -z "$order" ] && continue + if [ "$shown" -lt "$limit" ]; then + echo " ${order} ${id} ${acuity} — ${timing} · ${estimate}" + shown=$((shown + 1)) + else + more=$((more + 1)) + fi + done <`** — append a row to **Open items**. Infer `Pri`/`Type` from the text (ask only if genuinely ambiguous; default `P2`/`task`). Allocate the ID from the `` marker, then bump that marker. Fill `Source` with - `session ` unless the user names one; `Added` is today's date. + `session ` unless the user names one; `Added` is today's date. If the work is currently + recommended, also add it to the ordered execution ledger with acuity, intelligence, timing, + estimate, dependency, and completion signal; otherwise retain it only in Open items. - **`/issues done [outcome]`** — move that row from **Open items** to **Resolved / archive** - with today's date and a one-line outcome. Archive, never delete. + with today's date and a one-line outcome, remove it from the recommended execution queue, and + close the order gap. Archive, never delete. - **`/issues update `** — edit an open row's summary or next action in place. - **`/issues capture`** — scan the current session for recommendations, follow-ups, deferrals, and unfixed problems that surfaced but were not recorded. Propose them as a numbered list and add the @@ -54,6 +60,9 @@ paragraph; put the smallest next action in **Detail / next action**. - Keep the table format and column order exactly as in `docs/outstanding-issues.md`. One row per item. - IDs are monotonic and never reused — always allocate from the `issues:next-id` marker and bump it. +- Keep the recommended execution queue dependency-ordered, gap-free, deduplicated, and synchronized + with its referenced open rows. Never add refuted, parked, superseded, resolved, or decision-only + records to the active recommendation view. - Escape `|` inside cell text (write `\|`) so the markdown table stays intact. - Respect the repo's RAG/clinical/privacy flagging rules if an item _itself_ touches a protected surface — recording it here is fine, but acting on it later still needs the usual gate. diff --git a/AGENTS.md b/AGENTS.md index 32d9bc6c06..d8f0c38918 100644 --- a/AGENTS.md +++ b/AGENTS.md @@ -448,23 +448,29 @@ Run the matching planner command in `docs/productivity-workflows.md` without sid -## Outstanding-work memory (`/issues`) - -`docs/outstanding-issues.md` is the durable, cross-session memory of every outstanding **task**, -**recommendation**, and **issue** for this repo. Chat context resets between sessions; that file does -not, so anything worth remembering after a session ends belongs there. - +## Universal repository task ledger (`/issues`) + +`docs/outstanding-issues.md` is the single universal, durable, cross-session ledger for every +outstanding **task**, **recommendation**, and **issue** in this repo. It owns the recommended +execution order, acuity, required capability, timing, effort, approvals, evidence, status, success +criteria, stop rules, and resolution history. Chat context resets between sessions; that file does +not, so anything worth remembering after a session ends belongs there. Do not create or maintain a +second task ledger. + +- Before starting or recommending repository work, read the ordered **Recommended execution + queue** and the referenced open item. Only rows in that ordered view are active recommendations; + refuted, parked, superseded, resolved, and decision-only records remain audit history, not tasks. - When the user types `/issues`, invoke the `issues` skill (`.claude/skills/issues/SKILL.md`): read - `docs/outstanding-issues.md` and state the open items back, grouped by priority. A plain `/issues` - is read-only — it mutates and commits nothing. + `docs/outstanding-issues.md` and state the recommended execution queue in order, then summarize + the wider open-item counts. A plain `/issues` is read-only — it mutates and commits nothing. - `/issues add|done|update|capture …` mutate the ledger; each mutation commits **only** `docs/outstanding-issues.md` (no push unless the user asks or you are already handing off). - Proactively offer to `capture` unresolved follow-ups, deferrals, and known risks into the ledger before a session's context is lost — that is what keeps it a memory rather than a stale list. - A `SessionStart` hook (`.claude/hooks/issues-surface.sh`, wired in `.claude/settings.json`) - auto-surfaces the open items into context at the start of every session and, on a context reset - (`compact`/`resume`/`clear`), nudges a `/issues capture`. It is read-only — it never writes the - ledger. `/issues` is still the way to read the full list or mutate it. + auto-surfaces the ordered recommended tasks plus open-item counts at the start of every session + and, on a context reset (`compact`/`resume`/`clear`), nudges a `/issues capture`. It is read-only — + it never writes the ledger. `/issues` is still the way to read the full list or mutate it. ## Codex GitHub review behavior diff --git a/docs/README.md b/docs/README.md index 0b743e80cb..0d1c904af2 100644 --- a/docs/README.md +++ b/docs/README.md @@ -71,6 +71,7 @@ npm run docs:check-links ## Plans and workstreams (living) +- [outstanding-issues.md](outstanding-issues.md) — universal task ledger, recommended execution order, evidence, status, and resolution history - [maturity-backlog-workorders.md](maturity-backlog-workorders.md) — actionable work orders tracking the repository-maturity audit backlog - [framework-dependency-modernization-checklist.md](framework-dependency-modernization-checklist.md) — ordered Next.js 16, runtime, dependency, Turbopack, and verification migration program - [search-rag-master-plan.md](search-rag-master-plan.md) / [search-rag-master-context.md](search-rag-master-context.md) — search/RAG roadmap and shared context diff --git a/docs/branch-review-ledger.md b/docs/branch-review-ledger.md index b9b2a8b968..f8f2edb72f 100644 --- a/docs/branch-review-ledger.md +++ b/docs/branch-review-ledger.md @@ -20,7 +20,7 @@ Use this ledger to prevent repeated branch and PR reviews when the reviewed HEAD | Date | Branch or ref | Reviewed HEAD | Scope | Outcome | Checks | | ---------- | -------------------------------------------------------- | ---------------------------------------- | ---------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | -| 2026-07-24 | `codex/supabase-document-change-trigger` | `9c7d9edf509a51478f5bebbabcca64e3926dc877` + reviewed working diff | Document-change ingestion trigger migration, schema mirror, grants, privacy and fail-safe delivery | APPROVE. No P0-P2 finding. The trigger is update-only, acts solely on a strict JSON boolean false/absent-to-true transition, sends only the receiver's allowlisted owner-scoped fields, fails open for document writes when Vault/GUC/pg_net is unavailable, and revokes execution from public/anon/authenticated. No production URL fallback exists. Highest residual risk is deliberate pg_net at-most-once delivery; the clear-then-flip recovery and data-preserving rollback are documented, and the trigger remains inert until both the Vault secret and environment base-URL GUC are configured. | Disposable `supabase/postgres:17.6.1.127` schema replay and drift-manifest regeneration passed (16s; scratch container removed); focused schema/drift/receiver Vitest 89/89; migration-role, function-grant (30 SECURITY DEFINER functions) and owner-scope guards; production-readiness CI mode READY with expected secretless-worktree warnings; offline RAG 21 suites/307 tests; `verify:cheap` 365 files, 3,241 passed/1 skipped; static trace of receiver payload, authoritative owner-scoped reload and idempotent enqueue path. No live provider mutation or migration apply. | +| 2026-07-24 | `codex/supabase-document-change-trigger` | `9c7d9edf509a51478f5bebbabcca64e3926dc877` + reviewed working diff | Document-change ingestion trigger migration, schema mirror, grants, privacy and fail-safe delivery | APPROVE. No P0-P2 finding. The trigger is update-only, acts solely on a strict JSON boolean false/absent-to-true transition, sends only the receiver's allowlisted owner-scoped fields, fails open for document writes when Vault/GUC/pg_net is unavailable, and revokes execution from public/anon/authenticated. No production URL fallback exists. Highest residual risk is deliberate pg_net at-most-once delivery; the clear-then-flip recovery and data-preserving rollback are documented, and the trigger remains inert until both the Vault secret and environment base-URL GUC are configured. | Disposable Supabase Postgres 17.6.1.127 image replay and drift-manifest regeneration passed (16s; scratch container removed); focused schema/drift/receiver Vitest 89/89; migration-role, function-grant (30 SECURITY DEFINER functions) and owner-scope guards; production-readiness CI mode READY with expected secretless-worktree warnings; offline RAG 21 suites/307 tests; `verify:cheap` 365 files, 3,241 passed/1 skipped; static trace of receiver payload, authoritative owner-scoped reload and idempotent enqueue path. No live provider mutation or migration apply. | | 2026-07-23 | PR #1090 / `cursor/fix-phone-dock-edge-1b1d` | `761de7e9ad623b6bd8d634d849a9eb465d622e48` (merged as `09028ef217209fceb53f1122ac7738b509bce323`) | Phone safe-area and edge-to-edge search-dock UI review | MERGED. No P0-P2 finding. The branch was three commits behind, so current `origin/main` was merged before landing; the actual merge tree matched the reviewed synthetic tree. The dock remains flush to the viewport with safe-area padding inside the form, and the phone shell no longer retains the `dvh` clamp that created the Safari toolbar band. Zero actionable review threads. | `npm run ensure`; focused `ui-tools.spec.ts` phone-home and edge-to-edge scenarios: Chromium 2/2 and WebKit 2/2; refreshed hosted policy, security, unit, build, advisory UI, Production UI and required aggregate checks green; exact-head ancestry and local-main tree equality proved after merge. | | 2026-07-22 | PR #1087 / `codex/reconcile-product-truth` | `edbc2260fef59ca2fa7c6973dffb85e32354bce1` (merged as `05dc52fd8408a65117e22a6236e43252203bea92`) | Product-truth copy, account persistence and unavailable-SSO presentation | MERGED. Cross-device claims now match favourites/preferences persistence; recent searches are identified as browser-session data; the contradictory “never shared” statement is removed. All unavailable setup providers and Apple elsewhere use the connected accessible “coming soon” placeholder pattern. The single review finding was fixed, replied to and resolved. | Red DOM proof; focused 19/19; `verify:cheap` 3,220 passed / 1 skipped; `verify:ui` 265/265; PR-local build/secret scan/offline RAG; final hosted required, Production UI, policy and security checks green. No provider calls or RAG spend. | | 2026-07-22 | PR #1086 / `codex/reconcile-xlsx-budgets` | `5376880a40749b6526fd7e4603a7be9d04bc9624` (merged as `2963fba46eacd644618a588fa283f7597faa2644`) | XLSX resource-boundary review | MERGED. Enforces worksheet, non-empty-row, rendered-cell and UTF-8 output ceilings before result fragments are appended; sparse-column output is preserved. No actionable review threads. | Red 257-sheet reproducer; focused 4/4; `verify:cheap` 3,218 passed / 1 skipped; PR-local build/scan/offline RAG; hosted required/security/policy green. | diff --git a/docs/codebase-index.md b/docs/codebase-index.md index c7d597ed5a..aa92ad7ad2 100644 --- a/docs/codebase-index.md +++ b/docs/codebase-index.md @@ -334,6 +334,7 @@ One shared composer (`master-search-header.tsx`) serves every mode. Placement: | Full documentation index | `docs/README.md` | | Routes and modes | `docs/site-map.md` | | Search/RAG roadmap | `docs/search-rag-master-plan.md` | +| Universal task ledger | `docs/outstanding-issues.md` | | Reindex operations | `docs/reindex-runbook.md` | | Production readiness | `docs/production-readiness-checklist.md` | | Capacity / scale-up | `docs/capacity-review.md`, `docs/auth-connection-cap-runbook.md` | diff --git a/docs/maturity-backlog-workorders.md b/docs/maturity-backlog-workorders.md index 2c3a945c9d..3d207ed7f2 100644 --- a/docs/maturity-backlog-workorders.md +++ b/docs/maturity-backlog-workorders.md @@ -7,6 +7,10 @@ files**, **risk**, **verification**, and **status**. High-risk items are deliber their own work order — the audit's rule is one dedicated PR + full-suite verification per structural change, not a single mixed PR. +This file is supporting work-order detail, not a second task ledger. Only work represented in the +recommended queue in [`outstanding-issues.md`](outstanding-issues.md) is an active repository +recommendation; that universal ledger owns current status, priority, and execution order. + **Status legend:** `DONE` (landed) · `IN PROGRESS` (partially landed; more PRs remain) · `READY` (scoped, safe to start) · `OPEN` (needs a decision or a dedicated PR) · `PROVIDER-GATED` (touches live DB/CI/provider — needs explicit confirmation) · `SATISFIED` diff --git a/docs/operator-backlog.md b/docs/operator-backlog.md index eb749675b1..cc2933d296 100644 --- a/docs/operator-backlog.md +++ b/docs/operator-backlog.md @@ -1,8 +1,9 @@ # Operator backlog -Single source of truth for **human-only / provider-gated actions** that cannot be done from a coding +Detailed runbook index for **human-only / provider-gated actions** that cannot be done from a coding session (they touch Supabase, Railway, OpenAI, or GitHub settings, per the AGENTS.md provider boundary). -This exists so that launch-blocking state lives in the repo instead of chat memory. +Canonical task status, priority, and execution order live only in +[`outstanding-issues.md`](outstanding-issues.md); this file supplies operator procedures and evidence. **How to use:** work top to bottom; each row links to the detailed runbook. `Status` values are `⏳ pending`, `🔎 verify` (may already be done — confirm before repeating), `✅ done`, `—` (n/a). diff --git a/docs/outstanding-issues.md b/docs/outstanding-issues.md index e31b221520..f0bb324aea 100644 --- a/docs/outstanding-issues.md +++ b/docs/outstanding-issues.md @@ -1,4 +1,4 @@ -# Outstanding Issues, Recommendations & Tasks +# Universal Task Ledger — Outstanding Issues, Recommendations & Tasks Durable, cross-session memory of everything still outstanding for this repo: open **tasks**, **recommendations** not yet acted on, and **issues** not yet resolved. Chat context is ephemeral @@ -7,10 +7,14 @@ Durable, cross-session memory of everything still outstanding for this repo: ope **Rule of thumb:** if it is worth remembering after this session ends, it belongs here. +This is the repository's **single universal task ledger**. It owns recommended execution order, +acuity, required capability, timing, effort, dependencies, approvals, evidence, status, success +criteria, stop rules, and resolution history. Do not create or maintain a second task ledger. + ## How this is used -- Say `/issues` in Claude Code → the skill reads this file and states the open items back, - grouped by priority with a one-line summary count. Nothing is mutated on a plain read. +- Say `/issues` in Claude Code → the skill reads this file and states the recommended queue in + order, then summarizes the wider open-item counts. Nothing is mutated on a plain read. - `/issues add …`, `/issues done `, `/issues capture`, and friends mutate the tables below. The full command surface lives in the skill file. - Every mutation keeps this file committed so the memory survives across sessions and worktrees. @@ -27,7 +31,72 @@ Durable, cross-session memory of everything still outstanding for this repo: ope - Resolving an item moves its row to **Resolved / archive** with the date and a one-line outcome — rows are archived, not deleted, so the history stays auditable. - +## Execution scales + +- **A1 urgent:** active safety/privacy/data-loss/release blocker; `#057` and `#053` are the current A1 items. +- **A2 important:** confirmed correctness, clinical, privacy, reliability, or evaluation-integrity + work that leads its available lane. +- **A3 planned:** worthwhile work deferred to its stated trigger. +- **Optional:** do only when measured need, ownership, and cost justify it. +- **Standard:** experienced generalist. **High:** senior cross-module reasoning. **Specialist:** + database/RAG/clinical/privacy/security/evaluation expertise. **Operator:** authorised human owner. +- Effort is active work, excluding approvals, hosted/provider waits, soak, and review time. + +The order below is planning guidance, not authority to call providers, spend money, change +production, commit, push, merge, or deploy. A waiting dependency does not block an independent +executable item below it. + +## Recommended execution queue + +Last reconciled on **2026-07-24** against fetched `origin/main` +`1eed39ed2b43509634a8d745d75a6956edb60e83`. No live application or provider state was queried. +Revalidate the referenced evidence against current `main` immediately before starting a row. + +| Order | Source | Recommended outcome / next action | Acuity / capability | When | Active effort | Dependencies / approval | Done, verification, and stop rule | +| ----: | ------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | -------------------------------------------------- | ----------------------------------------------------------- | -------------------------------------- | -------------------------------------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | +| 1 | `#057` | Verify provider-side containment of credentials previously exposed outside authorised stores; revoke or rotate any still-valid OpenAI, Supabase service-role/database, and E2E credentials, then update only intended secret stores. | A1 / Operator security + independent reviewer | Immediate approved security window | 1–3 hours | Credentialed provider owners; correct target identity; explicit provider and secret-store approval | Every old credential is rejected or retired; replacements exist only in intended stores; presence/readiness checks pass; no value enters Git, logs, issues, or chat. Use approved provider evidence and secret scans. Stop before provider action without approval and do not rewrite history without separate evidence. | +| 2 | `#053` | Align the Safety Plan Generator with the repository's no-patient-data contract before real-patient use. | A1 / Specialist privacy + clinical | Ready; first | 3–5 hours | Privacy/clinical owner | Copy, behavior, print, and tests agree; no identifier is persisted/transmitted without an approved basis. Run focused UI/privacy/a11y checks, `verify:cheap`, production-readiness. Stop before adding storage/provider flow. | +| 3 | `#054` | Fail closed when answer relevance metadata is absent; add the red policy test, then make the smallest render-policy correction. | A2 / Specialist clinical safety | Ready; after `#053` | 2–4 hours | Local/offline | Missing and explicit-false relevance are conservative; explicit source-backed behavior is unchanged. Run focused policy/provenance tests, `verify:cheap`, production-readiness. Stop before retrieval/ranking/generation. | +| 4 | `#052` | Prevent full/retry reindex overlapping a fresh agent-enrichment lease; add red single/bulk tests, then reuse `hasActiveAgentEnrichmentJob` before mutation. | A2 / High concurrency | Ready; after `#054` | 0.5–1 day | Local/offline | Fresh processing leases conflict; stale leases and enrichment mode retain behavior. Run focused route/safety tests, `verify:cheap`, production-readiness. Stop if current `main` does not reproduce it. | +| 5 | `#055` | Recover aged `queued` documents that have no open ingestion job; begin with a stranded-row reproducer and the smallest idempotent owner-scoped recovery path. | A2 / Specialist queue reliability | Ready; after `#052` | 0.5–1.5 days | Local replay for schema/scheduling; hosted changes need approval | Exactly one recoverable job is created under crash/concurrency/repeat; fresh queues and open jobs are untouched. Run focused recovery/schema tests, disposable replay if needed, `verify:cheap`, production-readiness. Stop if safe age/ownership cannot be proven. | +| 6 | `#030` | Require distinct source identities for distinct expected comparison slots; begin with a red shared-title/two-expectation test. | A2 / High evaluation semantics | Decision-ready; code only after owner choice | 2–4 hours | Local/offline; protected-eval review | One source cannot satisfy both slots; existing cases stay green. Run focused matching, typecheck, `verify:cheap`. Stop before alias/retrieval/ranking changes. | +| 7 | `#019` | Prove admission evidence is lost at the comparison fallback boundary; scope the smallest correction without automatically shipping it. | A2 / Specialist RAG | Reproducer now; behavior after `#051`/`#023` | 0.5–1 day; fix separate | Protected review; approved baseline/post canary for behavior | Reproducer fails while packing retains both sources. Stop without behavior change if the boundary defect is not independently reproducible. | +| 8 | `#058` | Execute OpenAI/Railway DPAs; decide ZDR/residency; obtain cache/subprocessor answers and APP 8 plus APP 5/1 counsel sign-off. | A2 / Operator legal/privacy | Start now; before real-patient use/privacy-approved release | 4–8 hours internal; 1–6 weeks elapsed | Signers, counsel, providers | Record countersigned evidence, IDs, dates, decisions, and approved wording in PIA/cross-border docs. Stop before public-copy changes without counsel. | +| 9 | `#059` | Set a unique 32+ character `OPENAI_SAFETY_IDENTIFIER_SECRET` per environment with `OPENAI_API_KEY`, never in Git/output. | A2 / Standard local + Operator hosted | Local now; hosted approved window | 15–30 min local; 30–60 min/env | Secret manager; hosted approval | Presence-only readiness passes, HMAC pseudonymity remains, secret scans clean. Stop if governance rejects stable identifiers. | +| 10 | `#060` | Reconcile query-hash, deep-probe, Supabase/OpenAI, project identity, and schedules using read-only presence checks before setting anything. | A2 / Operator platform | Next approved readiness window | 1–2 hours | Provider approval, env owners | Record status, never values; set only confirmed gaps; identity/readiness pass. Stop on target ambiguity or cross-env key reuse. | +| 11 | `#022` | Decide auditable BMJ reference policy, then review the ten highest-impact local WA documents. | A2 / Operator governance + Specialist | Decision-ready | 1–2 hours policy; 0.5–1 day reviews | Clinical authority; live writes/canary approval | Preserve third-party/unverified provenance; record reviewer/evidence/time/rationale. Stop after ten and remeasure debt. | +| 12 | `#051` + `#023` | Compare run `30018289898` with the scheduled 2026-07-26 canary/browser/label artifacts; disposition residuals without rerun. | A2 / Specialist evaluation | About 2026-07-27 02:00 AWST plus runtime | 2–4 hours | Approval to read hosted artifacts; no dispatch | Record tree/run, content/provider/latency, browser, and labeling outcomes. Stop without spend or archived lithium changes. | +| 13 | `#018` | Diagnose lithium, ADHD, metabolic residuals independently: table fast path, extractive budget, schedule-free selection. | A2 / Specialist clinical RAG | After `#051`/`#023`; one mechanism at a time | 1–2 days diagnosis; fixes separate | Protected review; approved behavior canary | Each has a current-main fixture and scoped candidate; stop without deterministic repro or on individual canary regression. | +| 14 | `#001` | Reassess semantic reranking without changing default; enable only after accepted ambiguity comparison. | A2 / Specialist ranking | After `#051`/`#023` and rollout approval | 0.5–1 day plus canary review | Provider approval/budget | Keep off unless 36/36, recalls 1.0, zero case regressions, measured gain. Otherwise record keep-off decision. | +| 15 | `#033` | Decide whether governance metadata enters the LLM source block, distinguishing unknown from adverse status. | A3 / Specialist prompt governance | After `#022` and `#051`/`#023` | 1–2 days plus eval | Better metadata, stable canary, provider approval | Prompt/serialization tests; no grounded-supported drop; zero citation failures. Stop on over-caveating/degradation. | +| 16 | `#024` | Separate Playwright `_rsc` interception failure from a real Safari defect. | A3 / High Next.js/WebKit | After scheduled matrix data or real Safari repro | 0.5–1 day | Device/provider check only if local ambiguity | Test-only fix needs discriminating repro and meaningful access-control assertions; keep Chromium green. | +| 17 | `#007` | Choose canonical Tools experience; align nav, redirects, sitemap, reachability. | A3 / Operator product + Standard frontend | Product decision window | 15–30 min decision; 0.5 day code | Product owner | One canonical path and intentional redirect with green route/UI tests. Stop on unresolved standalone need. | +| 18 | `#025` + SLO alerts | Select owned deploy/CI/ingestion/SLO channels; configure only those, ingestion after `#055`. | A2 / Operator integrations | Approved observability window | 1–3 hours/channel | Provider approval, secrets, destination owner | Mocked tests then one controlled non-PHI event/channel. Stop on missing owner or unsafe delivery. | +| 19 | `#061` | Run release/clinical gates once against an exact candidate, not a moving branch. | A2 / High release + Operator | Before release/handoff needing full confidence | 0.5 day plus waits | Exact SHA, provider approval, heavy-command lock | All required gates recorded against one SHA; stop/classify first failure and do not repeat unchanged pass. | +| 20 | `#062` | Provision dedicated Clinical KB staging on Supabase/Railway with isolated keys and synthetic/non-clinical data. | A2 / Operator + Specialist DB | After cost/ownership approval | 0.5–1 day | Billable provider approval, owner | Identity, tenancy, secrets, migrations, app/worker health, data boundary pass. Stop on target ambiguity or production data/key reuse. | +| 21 | `#063` | Run documented soak and rollback rehearsal in dedicated staging. | A2 / Operator reliability | After staging provision | 0.5 day plus soak | Staging and provider approval | SLO, rollback, data integrity, recovery evidence pass. Stop before production if unproven. | +| 22 | `#064` | Verify registry/differentials/medications non-empty before writes; seed only confirmed governed gaps. | A2 / Operator data + Specialist | Approved production window | 1–3 hours plus seed time | Correct owner/project, provider approval | Read-only proof precedes idempotent writes; stop if populated or target ambiguous. | +| 23 | `#011` | Switch Auth DB cap to percentage immediately before first compute scale-up. | A3 / Operator capacity | Scale-up trigger only | 30–60 min plus soak | Dashboard, planned scale-up, approval | Target only `sjrfecxgysukkwxsowpy`; record before/after and advisor/health recheck. Do not create/use staging for this task. | +| 24 | `#017` | Capture reproducible mobile/desktop production LCP, INP, CLS before performance work. | A3 / High web performance | Before `#012`/`#013`; approved live window | 1–2 hours | Live-site approval; record route/build/throttle | Record metrics and accept/reject decision. Stop if acceptable or too noisy. | +| 25 | `#037` | Decide whether routine supported claims cap at medium trust. | A3 / Operator clinical-product + Standard frontend | Next trust-policy review | 30–60 min decision; up to 0.5 day code | Clinical/product authority | Record policy; if accepted, change flag/render expectations only and run focused tests. Stop without owner acceptance. | +| 26 | `#027` | Add uptime monitor outside GitHub/Railway only if service, budget, responder justified. | Optional / Operator SRE | Owned external-alert trigger | 1–2 hours | Vendor/privacy/owner/provider approval | Controlled non-PHI failure/recovery alerts. Stop without responder or if current monitoring accepted. | +| 27 | `#028` | Define vendor/region/retention/redaction/sampling/maps/owner before runtime error tracking. | Optional / Specialist privacy + Operator | After privacy/ownership/cost approval | 1–3 days | Privacy and provider approval | Redaction tests exclude clinical data/IDs/secrets; one non-sensitive event arrives. Stop if envelope unacceptable. | +| 28 | `#012` + `#013` | Re-measure route payloads and slim only one production route with a demonstrated budget problem; ignore mockup-only code unless it enters production. | A3 / High Next.js performance | After `#017` or equivalent evidence; one route | 0.5–2 days/route | Measured target, Next.js guide, UI verification | Demonstrate a material parsed/gzip or interaction gain with unchanged behavior; run analysis, focused tests, `verify:cheap`, browser smoke. Stop if gain is small. | +| 29 | `#040` | Add small stable visual-regression baselines with owner/update workflow. | Optional / High visual QA | Stable surfaces and owner | 1–2 days | Stable browser, owner | Low-flake repeat runs and documented updates. Stop before blocking if churn high. | +| 30 | `#038` | Define shared comparison interaction contract before another comparison surface; keep clinical content mode-specific. | Optional / High design-system | Approved new comparison surface | 0.5–1 day | Concrete surface, product/design owner | Inventory patterns and make shared behavior testable. Stop if no new surface. | +| 31 | `#035` | Expand conflict detection only for concrete clinically reviewed class with positive/negative fixtures. | A3 / Specialist evidence rules | Demonstrated missed conflict | 0.5–1 day design; code separate | Clinical review; provider approval only for live validation | Fixtures discriminate class without unrelated warnings. Stop if no bounded class. | +| 32 | `#039` | Converge repeated catalogue toolbar behavior only during a concrete toolbar project. | Optional / High frontend architecture | Concrete project | 0.5–1 day inventory; 1–3 days code | Product/design owner, UI verification | Prove shared behavior without flattening search semantics; stop after bounded contract. | +| 33 | `#056` | Write a product/privacy/persistence brief for “Current Clinical Work” before any storage or UI implementation. | A3 / High product architecture + privacy | Only when the product owner wants to evaluate the feature | 0.5–1 day | Product owner, privacy/retention decision, evidence of demand | Define users, data classes, lifecycle, cross-device expectations, deletion, failure states, and a smallest testable slice. Stop if demand or safe persistence cannot be established. | + +### Queue maintenance + +1. Revalidate against current `main` before starting and remove/rewrite contradicted work. +2. Do not combine protected RAG residuals or change scores, comparators, aliases, clamps, or semantic + reranking without separate reproducers and required validation. +3. Treat provider status as a claim: verify only after approval and never store secret values. +4. Close/reclassify when a success or stop condition is met; do not preserve work for its own sake. + + ## Open items @@ -37,8 +106,21 @@ Durable, cross-session memory of everything still outstanding for this repo: ope | ID | Pri | Type | Summary | Detail / next action | Source | Added | | ---- | --- | ----- | --------------------------------------------------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | --------------------------------------------------------------------------------------------------------------------------------------- | ---------- | +| #058 | P1 | task | Close the APP 8, DPA and ZDR governance position | Execute the documented OpenAI/Railway legal and privacy package: record signed DPA versions, ZDR/residency and prompt-cache/subprocessor decisions, accountable owners and approved APP 8 plus APP 5/1 wording. Stop before real-patient use or privacy claims while the basis is unresolved. | `docs/openai-cross-border-basis.md`; `docs/privacy-impact-assessment.md` | 2026-07-24 | +| #059 | P2 | task | Configure the OpenAI safety-identifier secret safely | Set a unique 32+ character `OPENAI_SAFETY_IDENTIFIER_SECRET` in each approved environment that has `OPENAI_API_KEY`, using only its secret manager. Verify presence and HMAC pseudonymity without exposing values; stop if governance rejects stable identifiers. | production readiness checks; session 2026-07-24 | 2026-07-24 | +| #060 | P2 | task | Reconcile provider configuration without exposing values | In an approved readiness window, verify presence/status for query-hash, deep-probe, Supabase/OpenAI, project-identity and schedule settings before changing anything. Record status only, set confirmed gaps separately, and stop on target ambiguity or cross-environment key reuse. | `docs/operator-backlog.md`; session 2026-07-24 | 2026-07-24 | +| #061 | P2 | task | Run one full gate against an exact release candidate | Before a release or handoff needing full confidence, record one exact SHA and run the required release and clinical-quality gates once under explicit provider/spend approval. Stop and classify the first failure; do not repeat an unchanged pass. | `docs/production-readiness-checklist.md`; session 2026-07-24 | 2026-07-24 | +| #062 | P2 | task | Provision isolated Clinical KB staging | After cost and ownership approval, provision dedicated Supabase/Railway staging with isolated credentials and synthetic or non-clinical data. Verify identity, tenancy, migrations and app/worker health; stop on target ambiguity or production data/key reuse. | `docs/staging-setup.md`; session 2026-07-24 | 2026-07-24 | +| #063 | P2 | task | Prove staging soak and rollback readiness | After #062, execute the documented staging soak and rollback rehearsal under provider approval. Record SLO, rollback, data-integrity and recovery evidence; stop before production if any required recovery proof is missing. | `docs/launch-operator-runbook.md`; `docs/capacity-review.md` | 2026-07-24 | +| #064 | P2 | task | Verify production catalogue data before any seed | In the approved production window, verify registry, differentials and medications are non-empty before writes. Seed only a confirmed governed gap with correct owner/project identity and rollback; stop with no mutation when valid data already exists or identity is unclear. | `docs/launch-operator-runbook.md`; session 2026-07-24 | 2026-07-24 | +| #057 | P1 | task | Verify containment of previously exposed credentials | **Outcome:** credentials previously exposed outside authorised stores can no longer authenticate. **Next:** in approved provider dashboards, revoke or rotate any still-valid OpenAI, Supabase service-role/database, and E2E credentials; update only intended secret stores and record status without values. **Success:** provider evidence shows every old credential rejected or retired, replacements are scoped to intended environments, and readiness checks pass. **Verify:** approved provider audit/rotation evidence, presence-only checks, and secret scanning that never prints values. **Stop:** no provider or secret-store action without explicit approval; never paste values into Git, logs, issues, or chat, and do not rewrite history without separate evidence. | session 2026-07-24 security reconciliation | 2026-07-24 | | #051 | P2 | task | Stabilise the live answer-quality canary before more RAG tuning | Diagnostics landed in PR #1095: structured JSON/Markdown artifacts now record the actual checked-out SHA, run identity and latency context, and the offline trend tool separates content, provider-route and latency outcomes. First validating run `30018289898` recorded the expected tree and cost, with 36/36 retrieval green, but one report cannot establish variability; PR #1097 prevents a single failure being mislabeled as repeated. Next: compare the scheduled 2026-07-26 structured report with this run. Do not spend on an immediate retry or reapply the archived lithium guard before that comparison. | PR #1095; run `30018289898`; PR #1097; archive ref `refs/archive/rejected-rag/20260723/monitoring-subject-gate` | 2026-07-23 | | #001 | P2 | task | Semantic reranking still gated off | `RAG_SEMANTIC_RERANK_ENABLED=false` from PR #901. Do not enable until the provider-backed 36/36 retrieval-quality gate **and** an ambiguity-focused canary are explicitly approved and recorded. | `docs/process-hardening.md` (Semantic reranking rollout debt); PR #901 | 2026-07-21 | +| #052 | P2 | issue | Reindex can overlap a fresh agent-enrichment pass | Full/retry reindex preflights check `ingestion_jobs` but do not call the existing `hasActiveAgentEnrichmentJob`. A fresh `indexing_v3_agent_jobs.status='processing'` lease can therefore overlap destructive artifact work. Extend single and bulk full/retry preflights before mutation; keep stale leases non-blocking and preserve enrichment-mode behavior. | `src/lib/ingestion-mutation-safety.ts:116-159`; single/bulk reindex routes; ingestion audit 2026-07-24 | 2026-07-24 | +| #053 | P1 | issue | Safety Plan Generator contradicts the privacy contract | **Outcome:** the tool, privacy notice, PIA, and tests agree on whether patient identifiers may be entered, copied, printed, or saved. `patient-safety-plan.tsx` asks for “Patient (name or initials)” and produces a patient copy, while `/privacy` and the PIA say the product does not ask for patient data. **Next:** obtain a privacy/clinical decision, then default to identifier-free behavior unless transient identifier processing is explicitly approved and documented. **Success:** no contradictory copy; no identifier is persisted or transmitted without an approved basis. **Verify:** focused component/privacy/copy/accessibility tests, browser print/copy smoke, `verify:cheap`, production-readiness. **Stop:** any new storage or provider transmission requires a separate review. | `src/components/patient-safety-plan.tsx:649`; `src/app/privacy/page.tsx:28`; `docs/privacy-impact-assessment.md:18` | 2026-07-24 | +| #054 | P2 | issue | Missing answer relevance metadata is treated as source-backed | **Outcome:** absent `relevance` metadata renders conservatively. `RagAnswer.relevance` is optional, but `relevance?.isSourceBacked !== false` treats `undefined` as source-backed. **Next:** add the red render-policy test, then make the smallest policy-only fix. **Success:** missing and explicit-false relevance fail closed; explicit source-backed relevance is unchanged. **Verify:** focused answer-render-policy, provenance, clinical-safety, `verify:cheap`, production-readiness. **Stop:** do not expand into retrieval, ranking, or generation. | `src/lib/types.ts:1029`; `src/lib/answer-render-policy.ts:145` | 2026-07-24 | +| #055 | P2 | issue | Upload crash can strand a queued document without a job | **Outcome:** a crash between document and job creation cannot strand an upload indefinitely. **Next:** add the stranded-row reproducer, then choose the smallest idempotent atomic-enqueue RPC or bounded scheduled sweep consistent with current ownership and rollback contracts. **Success:** exactly one recoverable job is created; existing open jobs do not duplicate; owner scope, retry, audit, and rollback remain intact. **Verify:** focused upload/recovery/schema tests, migration guards, disposable replay if needed, `verify:cheap`, production-readiness. **Stop:** hosted changes require approval, and no at-least-once claim is valid until the crash case passes. | `src/app/api/upload/route.ts`; `docs/webhooks.md:163-192` | 2026-07-24 | +| #056 | P3 | rec | Define “Current Clinical Work” before implementation | **Outcome:** decide whether a workspace combining saved comparisons, partial formulation work, recent tools, and pinned source sets is worth building. **Next:** write a product/privacy/persistence brief only; do not build storage or UI. **Success:** the brief defines users, data classes, lifecycle, cross-device expectations, deletion, failure states, evidence of demand, and the smallest testable slice. **Stop:** close the idea if demand or safe persistence cannot be established. | session 2026-07-24; `src/lib/tools-catalog.ts:253-262` | 2026-07-24 | | #005 | P3 | rec | `finalScore` saturates at clamp ceiling | Base + ~40 stacked boosts routinely exceed 1.0, so strong matches tie at 1.0 and order by an arbitrary `document_id` tiebreak. If ranking is ever revisited, break ties by the **pre-clamp** score rather than raising the `[0,1]` ceiling (downstream gates assume `[0,1]`). Ordering already sorts by the unbounded pre-clamp `rankScore` (`clinical-search.ts:1735,1927,1950-1955`), so the clamp confines only the reported confidence value, not result order. Not a defect on the current golden set; any change here is a protected RAG surface (canary required). | `docs/rag-hybrid-findings-and-todo.md` P1 item 4; `src/lib/clinical-search.ts:1735` | 2026-07-21 | | #007 | P3 | rec | `/tools` vs `/?mode=tools` parallel Tools entry points | `/tools` (standalone `ApplicationsLauncherPage`) has no inbound in-app link; the sidebar Tools item uses `/?mode=tools`. Decide the canonical entry point and wire nav consistently, or drop the standalone `/tools` page + `/applications` redirect. Currently allowlisted in `tests/route-reachability.test.ts`. | `src/app/tools/page.tsx`; `src/app/applications/route.ts` | 2026-07-21 | | #009 | P3 | rec | Confirm `/api/jobs` is intentionally server/ops-only | No client `fetch()` reaches `/api/jobs` (only tests import it). Confirm it is a deliberate ops/manual surface; if abandoned, remove it. | `src/app/api/jobs/route.ts` | 2026-07-21 | @@ -46,7 +128,6 @@ Durable, cross-session memory of everything still outstanding for this repo: ope | #011 | P3 | task | Auth DB-connection allocation is operator-only | Supabase Auth (GoTrue) is capped at ~10 absolute DB connections (Supabase perf advisor). Switch to **percentage-based** allocation in the Supabase **dashboard** before the first compute scale-up — **not settable via SQL/MCP** (operator-owned). Verify via a staging soak + an approval-gated read-only advisor re-check. | `docs/auth-connection-cap-runbook.md`; `docs/process-hardening.md` (Known follow-up debts) | 2026-07-21 | | #012 | P3 | rec | Slim the lazy cross-mode differentials chunk | `cross-mode-differentials.ts` is dynamically imported (correctly code-split **out** of the initial/dashboard bundle — verified), but it pulls the full ~860 KB differentials snapshot (~125 KB gzip lazy chunk) just to build a tiny `{slug,title,clinicalHinge}` + presentations + aliases catalog. A precomputed lightweight index (generator + drift check, like the `specifiers-content` split / medications `fields=index`) would cut that lazy chunk ~5–10×. Not a bundle leak — an M-effort slim. | `src/lib/cross-mode-differentials.ts`; `src/components/clinical-dashboard/cross-mode-links.tsx:150`; session 2026-07-21 (build:analyze) | 2026-07-21 | | #013 | P3 | rec | Route-chunk + mockup catalogue JSON weight | `build:analyze`: `/specifiers` ships `specifiers-search-index.json` (~180 KB parsed), `/forms` ships `forms-catalog.json` (~132 KB), `/formulation` ships `formulation-content.json` (~52 KB, client-side local search — needs index/full split or a search endpoint, architectural). All route-scoped (not initial bundle). Also `*-mockups.tsx` (~100 KB across chunks) build though `/mockups` 404s in prod — exclude from the prod artifact. | session 2026-07-21 (build:analyze) | 2026-07-21 | -| #014 | P3 | rec | Realize the `next/image` win on signed previews | `next.config` `images` (AVIF + `*.supabase.co` `remotePatterns` pinned to the project host, from #1024) is currently inert — signed document/image previews still render as raw ``. Route them through `next/image` to actually get AVIF + lazy optimization. | #1024; `src/components/clinical-dashboard/signed-image.tsx`; session 2026-07-21 | 2026-07-21 | | #016 | P3 | rec | "Big but not easy" structural + motion perf | Deferred larger levers: (a) nonce-CSP forces every product route to `ƒ Dynamic` (zero static generation) — evaluate Partial Prerendering / static shells for the static clinical catalogues (DSM/differentials/therapy/specifiers/formulation); (b) sidebar expand/collapse animates `grid-template-columns` (biggest smoothness cost, motion-gated — needs a transform-overlay rethink); (c) Therapy Compass fetches 692 KB / 2.5 MB JSON client-side (defer until interaction + confirm brotli); (d) settings/setup/admin dialogs static-imported into the home chunk (`next/dynamic` them). | session 2026-07-21 (build route table + design audit) | 2026-07-21 | | #017 | P3 | task | Field Web-Vitals baseline via live Lighthouse | In-sandbox runtime vitals were blocked (prod server hard-requires Supabase secrets; dev-mode CLS measured excellent at 0.00–0.04, content-first pages 0.000). Run Lighthouse against `psychiatry.tools` for real LCP/INP/CLS to prioritize #012–#016 by measured impact rather than reasoning. | session 2026-07-21 (measurement pass) | 2026-07-21 | | #018 | P2 | task | Split the lithium, ADHD and metabolic residuals by mechanism | Revalidated on current main 2026-07-23: these are not one composer defect. Lithium reproduced an unrelated-table retrieval fast-path defect; ADHD retrieves a relevant chart-heavy CAMHS source but exhausts the extractive route budget; metabolic retrieves the correct AKG source but selects schedule-free prose. The narrow lithium subject-evidence guard improved targeting from 0 to 1 with golden recall 1.0 and no reciprocal-rank regressions, but it was reverted because the full canary failed. After #051 stabilises the canary, add independent current-main reproducers and assess each mechanism separately. Do not widen the matcher or combine these into a broad ranking/composer change. | runs `30007833352` and `30009207429`; PR #1093; session 2026-07-23 | 2026-07-21 | @@ -62,7 +143,6 @@ Durable, cross-session memory of everything still outstanding for this repo: ope | #030 | P3 | issue | Wide-tier alias lets one doc satisfy both comparison slots | In src/lib/eval-document-matching.ts, "Admission to Discharge for Mental Health Inpatients" appears in BOTH the AdmissionCommunityPts and Discharge alias lists, so a single document can satisfy both expectedFiles slots and make allHit true — a latent false-pass on admission-discharge cases. Not firing today (that doc is not in the failing top-5) but it would mask a real miss. Tighten the tables so one doc cannot fill both sides. | src/lib/eval-document-matching.ts:32-65; session 2026-07-22 | 2026-07-22 | | #032 | P3 | rec | Governance ranking weighting: REFUTED, not debt | The source-governance audit (PR #1051) flagged three "gaps": `review_due` carries no ranking penalty, `unknownCurrentnessPenalty` ships at 0, and `selectBestSourceRecommendation` ignores governance metadata. **These are deliberate, measured decisions — do NOT implement them as written.** Blanket metadata boosts/penalties in selection ordering were measured on 2026-07-02 to regress the golden retrieval eval to 16/23 (doc-recall@5 1.0→0.76, mrr 0.75→0.64). Two corpus facts make it unsafe: scores saturate at the clamp so stacked boosts fully override lexical relevance, and the corpus is only partially metadata-enriched while `normalizeSourceMetadata` coerces unenriched docs to `unknown`/`unverified` — so "unknown" ≠ "bad" and blanket weighting swings ranking approx. 0.35 for reasons unrelated to relevance. Even governance-as-tiebreak buried correct unenriched docs (3 designs bisected). Next action: none — treat as a guardrail. If ever revisited, RC8 (source-strength as a _filter_) is the tracked path, gated on `eval:retrieval:quality` 36/36 plus a live canary pair. | PR #118; `docs/rag-behaviour/refuted-approaches.md`; PR #1051 items 4/5/6 | 2026-07-22 | | #033 | P3 | rec | Source governance metadata absent from the LLM prompt | `buildRagSourceBlock` omits `document_status`, `clinical_validation_status`, and `extraction_quality`, so the model cannot self-caveat during generation and governance is enforced only post-hoc. Generation-surface change: needs `eval:rag` plus `eval:quality --rag-only` (grounded-supported must not drop, citation-failure 0) and explicit approval. Carries the same "unknown ≠ bad" hazard as #032 — on a partially-enriched corpus the model would likely over-caveat correct sources, so design the prompt wording before spending an eval. | `src/lib/rag/rag-source-block.ts:126-198`; PR #1051 audit item 8 | 2026-07-22 | -| #034 | P3 | issue | Answer cache can serve stale governance metadata | `cacheIndexingVersion` derives the version from `updated_at` / `indexed_at` / `index_generation_id`, so a metadata-only `document_status` flip that bumps none of those is invisible to the passive guard. **Already mitigated**: every known status-write path calls `invalidateRagCachesForOwner` or `invalidateRagCachesForDocumentMutation`. Residual risk only — a future write path that omits the invalidator would serve stale governance until TTL. Next action: add a regression test pinning the invalidator call on status-mutating routes (cheaper and safer than touching the protected cache key). | `src/lib/rag/rag-cache.ts:382-438`; PR #1051 audit item 10 | 2026-07-22 | | #035 | P3 | rec | Threshold-conflict detection covers only 3 params | `detectThresholdDisagreements` checks only ANC, WBC, and platelets paired with withholding verbs, so cross-source conflicts on medication doses, lithium/thyroid levels, or vital signs go undetected. Deliberately narrow (see the comment at `:469-474`). Broadening changes when an answer is classified `conflicting` and adds warnings — real false-positive risk. Needs new fixtures plus a behaviour review before any change. | `src/lib/evidence.ts:469-574`; PR #1051 audit item 7 | 2026-07-22 | | #036 | P3 | rec | No explicit `is_public` visibility flag on documents | Public-corpus visibility is implicit: `owner_id IS NULL` on an `indexed` document (`resolveSearchScope`). The `metadata.public_corpus` marker is written by the promotion migrations but never used as a retrieval filter. Promotion is unconditional on `clinical_validation_status`, so unverified documents are publicly searchable — compensated by keeping `unverified_source` in the frontend-visible warning set. A hard schema flag touches RLS and the clinical-risk-gated retrieval RPCs; weigh against the existing compensating control before acting. | `supabase/schema.sql:61-108`; `src/lib/search-scope.ts:181-236`; PR #1051 audit item 3 | 2026-07-22 | | #037 | P3 | rec | D5 trust-cap-all-claims flag parked OFF | `NEXT_PUBLIC_RAG_TRUST_CAP_ALL_CLAIMS` extends authority gating from high-risk claims to **all** supported claims (`deriveTrust`). Ships OFF by design; flipping it caps trust to `medium` for routine claims across the board — a product/clinical-UX decision, not a defect. Both states are test-pinned. Next action: product decision, then flip and re-baseline the UI expectations. | `src/lib/answer-render-policy.ts:159-177`; PR #1051 audit item 11 | 2026-07-22 | @@ -78,6 +158,8 @@ Move resolved rows here with the resolution date and a one-line outcome. Keep th | ID | Type | Summary | Outcome | Resolved | | ---- | ----- | ---------------------------------------------------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ---------- | | #026 | task | Wire the Supabase document-change trigger | PR #1100 merged after disposable PostgreSQL replay and hosted migration replay. Production migration history and read-only catalog proof confirm the enabled metadata trigger, security-definer function, pinned search path and denied anonymous/authenticated execution; `npm run check:drift` reports no unexpected live drift. Delivery remains intentionally inert until the operator inputs tracked in #025 are configured. | 2026-07-24 | +| #034 | issue | Answer cache can serve stale governance metadata | Current-source verification found direct route coverage already asserts RAG-cache invalidation on document PATCH, source review, label, bulk and reindex mutation paths. The residual test recommendation is therefore already met; changing the protected cache key is unnecessary. | 2026-07-24 | +| #014 | rec | Realize the `next/image` win on signed previews | Superseded: `SignedImage` uses `next/image` for layout/sizing but deliberately sets `unoptimized`, preventing bearer signed URLs from entering the unauthenticated optimizer cache where cached content could outlive the token. No optimization task remains unless private-image delivery changes. | 2026-07-24 | | #031 | issue | Populate canary Source Governance table | The answer-quality step now consumes the preceding `golden-retrieval.json` only for source-governance reporting. Offline replay of run `30018289898` populated 338 top results, including 202 review-required entries, while retaining zero retrieval cases and no additional threshold failures. Retrieval and ranking behavior are unchanged. | 2026-07-24 | | #020 | task | Validate eval:quality cost readout post-fix | Confirmed on merged-main canary run `30018289898`: Answer Metrics reported 9 nonzero-cost cases and an estimated answer cost of `$0.234736`; the structured report retained the same value. The PR #1050 estimator fix is operationally proven. | 2026-07-23 | | #003 | task | Staging tenancy release evidence outstanding | Ran GitHub Action and validated isolation | 2026-07-21 |