Skip to content
Merged
2 changes: 2 additions & 0 deletions docs/branch-review-ledger.md
Original file line number Diff line number Diff line change
Expand Up @@ -923,8 +923,10 @@ Records before 2026-07-28 were written by hand and had drifted: 146 lines carrie
| 2026-08-12 | codex/implement-process-safety-for-multi-agent-workflows | 6d054c1fa02a988829def3274b32d31c13570851 | full PR diff and unresolved review feedback | Fixed cached-origin truthfulness, agent-safe approval gates, UI browser proof ordering, index handoff safety, and synced main | focused Vitest 31/31; tsc --noEmit pass; git status clean |
| 2026-08-13 | PR-1845 | 795ce38e165ce44e167038839a914b2efdb77dae | current-main merge, CI repair, and open-comment review | preserved the current-main ledger; canonicalized whitespace-only q to the non-empty legacy query; replaced the stale clear-filter Playwright locator; no unresolved review threads remained | pending fresh GitHub CI |
| 2026-08-12 | PR #1854 / codex/chat-differentials-results-design-differentials-results-design | 86f7d22c0b5d6204708547740fc44228853a9662 | review-and-fix | Fixed the append-only ledger conflict and added a truthful zero-count result-type empty state with reset action; synced current main. | focused Differentials DOM test passed; typecheck passed; fresh hosted CI required on final head |
| 2026-08-12 | claude/design-issues-triage-wnr7k9 | e6eb4c64b28141654bf5eb38121e657f3c2ef77c | Capture two session findings into the ledger (#310 fuzzy drug-match, #311 loss-detector) | docs-only; #310 records a measured cross-drug fuzzy match in open PR #1851 (fluoxetine->duloxetine at distance 2) with a tested one-line cap; #311 records promoting the derived ledger loss-detector to scripts/ | verify:pr-local 10/10 green after npm ci; check:outstanding-issues 114 open/195 archived |
| 2026-08-12 | codex/codex-cloud-github-action-bridge | c633b92ed5cecac6495a2292ff7b84b11348120d | credential-isolated Codex Run PR operator security and reliability review | 3 findings fixed: trusted base-merge accounting, fail-closed dispatch and mutation preflight, generated-comment and rerun verification hardening | npm run test:ci-workflows (274 passed, 11 skipped); npm run check:github-actions; YAML parse plus 8 bash and 2 github-script syntax checks; clean merge of origin/main |
| 2026-08-12 | origin/pr/1850 | 385a0795d6c6f36e69c60d3b5424873115ea3e99 | PR #1850 full diff vs origin/main | P2 and CI focus regression fixed | focused DOM and Chromium pending coordinator; prior CI static build and UI passed |
| 2026-08-12 | 1850 | a08a0f6794da6990aae0d3446c43eb37f51b7f84 | PR #1850 full diff vs origin/main | merge-conflict resolved cleanly; no remaining actionable findings | merge-tree clean; installed-lock-parity pass; focused in-page-nav DOM 34/34 pass (single worker); changed-file format pass; pre-merge audit pass; hosted CI pending |
| 2026-08-12 | codex/specifiers-results-polish-20260813 | 3e82cca69a72a66c1be87c5b9357336c3a95b7b0 | specifier result-card layout and interaction | No findings after resolving reduced-motion, dark-mode, and focus-ring review items | focused Chromium 1/1; lint pass; typecheck pass; RAG fixtures 36/36; full unit suite has 17 unrelated Windows/tooling baseline failures |
| 2026-08-12 | codex/specifiers-results-polish-20260813 | bbdb8337c2784941664a88a0d35a96a8c96a2edb | specifier result-card layout and interaction | Current-main sync introduced no changes to reviewed Specifiers scope; no findings after resolved review items | focused Chromium 1/1; lint pass; typecheck pass; RAG fixtures 36/36; full unit suite has 17 unrelated Windows/tooling baseline failures |
| 2026-08-12 | PR #1870 / claude/design-issues-triage-wnr7k9 | 6edfeceeb768b5714f98dad355361ad9148374b0 | review-and-fix | Merged current main; corrected #310's per-record fuzzy-trigger analysis and regression-test condition; preserved #311; removed the temporary self-mutating workflow; no additional P0-P2 findings in a distinct adversarial pass. | verify:pr-local -- --files docs/branch-review-ledger.md,docs/outstanding-issues.md; check:outstanding-issues; check:branch-review-ledger; exact-head hosted CI pending |
4 changes: 3 additions & 1 deletion docs/outstanding-issues.md
Original file line number Diff line number Diff line change
Expand Up @@ -137,7 +137,7 @@ removed after current-main verification; it is not missing recommended work.
| 82 | `#257` | Optional | High — formulation/specifiers flake | Standing until second reproduction | 15–30 min | Single unreproduced ui-formulation flake when run with ui-specifiers — record a second sighting only; do not quarantine until three on the same SHA. **Stop:** do not weaken assertions. |


<!-- issues:next-id=310 -->
<!-- issues:next-id=312 -->
## Open items

> **Merged-main canary update (2026-07-23, run `30018289898`):** the new structured report correctly recorded evaluated tree `c24f2e8f2d30d0c59fc1eba025d3dcd63478137e`, run/attempt identity and `cross-region-runner` latency context. Golden retrieval remained 36/36 with document/content recall 1.0 and no failed cases. The 44-case answer gate had grounded-supported and unsupported-correct rates of 1.0, but failed because `neuroleptic-side-effect-escalation` again returned one citation where two are required (citation-failure rate 0.0227). `admission-discharge-comparison` again omitted the specific AKG admission document after `comparison_source_extractive_fallback`; `admission-discharge-coverage-paraphrase` was advisory-only at 24,870 ms. Answer cost was reported as `$0.234736`. Do not retry immediately: retain this as the first structured datapoint, compare it with the scheduled 2026-07-26 report, and keep retrieval/ranking unchanged.
Expand Down Expand Up @@ -267,6 +267,8 @@ removed after current-main verification; it is not missing recommended work.
| #305 | P3 | rec | Canary has no latency-mode coverage and its cost readout is a known lower bound | Two informational gaps from the 2026-08-12 canary review, deferred by scope decision. (1) eval:retrieval:latency (p90 20s gate) is never wired into eval-canary.yml, so live retrieval latency regressions are invisible to the weekly canary while the answer step relaxes its own gates via EVAL_LATENCY_CONTEXT=cross-region-runner. (2) estimated_cost_usd applies one rate set (gpt-5.6-terra) to all usage including 2x-priced strong-model retries, so any cost trend understates strong-retry runs — the workflow comments say so, but eval:trend consumers may not read them. Also noted: the workflow-wide concurrency group (eval-canary, cancel-in-progress false) can queue a dispatched pair run behind a scheduled run, interleaving pair evidence; and fixture coverage gaps tracked in #018 remain uncatchable by the canary. Next: decide whether a monthly latency-mode dispatch is worth the spend; add a strong-usage split to the estimator if cost trends start driving decisions. | session 2026-08-12 RAG canary review | 2026-08-12 |
| #308 | P3 | issue | Desktop /documents/search CLS is 0.119, above threshold and stable across runs and baselines | Measured 2026-08-12 during the #147 close-out, twice, on the offline Lighthouse harness (Chromium 141): desktop /documents/search CLS **0.119**, against a committed baseline that also reads **0.119**. So this is long-standing and deterministic, not a regression — and it is above the 0.1 threshold. It sits outside #147's scope, which was mobile only, and it contradicts that row's claim that 'desktop passes everywhere: 0.016-0.097' — that range is stale. Companion desktop values from the same runs, all passing: /dsm 0.014, /forms 0.059-0.064, / 0.006, /therapy-compass 0.000. Next: attribute it the way #147 was attributed — drive Chromium against the offline production build with a PerformanceObserver on layout-shift reading entry.sources[].node, at DESKTOP emulation this time. Do not assume it is the same phone-overlay reserve cause as #147; that reserve publishes 0px above the phone breakpoint by construction, so this is a different shifter. Stop: do not raise the budget to accommodate it, and do not read local LCP or TBT from that harness (loopback has no network latency). | Local offline verify:lighthouse runs 2026-08-12 (two runs, identical CLS); #147 close-out; lighthouse-budget.json | 2026-08-12 |
| #309 | P2 | task | Facet groups of 6-20 options render as chips, not the dense list docs/filter-contract.md section 5 requires | Raised by the Codex reviewer on PR #1858 (P2) and correct. Section 5 of docs/filter-contract.md sets density by option count: <=5 chips, 6-20 dense full-width list with a right-aligned count column and group headings, >20 or >3 groups adds find-a-filter and collapse-by-default. Formulation is the first real facet adoption and renders NINE derived domains, so it should be in the 6-20 band, but ResultFilterFacetChips only has one layout — a wrapping chip row — and it is used at both breakpoints. So the first adoption sets the precedent of bypassing the density rule. **Why it was not fixed in #1858:** making the shared renderer density-aware changes a component every facet consumer will use, needs a layout decision at both breakpoints, and overlaps the >20 tier that PR F is scheduled to port up from the documents panel (which already has its own dense list, find-a-filter and collapse). That is a design decision and a PR of its own, not a scoped fix inside a mode adoption. **Next:** decide whether the shared renderer grows a density prop derived from options.length, or whether PR F's port-up supplies the dense tier and facet modes adopt it then; either way add a nine-option DOM assertion so the band is pinned rather than incidental. **Stop:** do not add a per-mode dense list — a second hand-rolled facet layout is exactly the breakpoint drift the shared renderer was extracted to remove. Services (PR C) will have six groups and hits the >3-groups branch of the same rule, so this should be settled before or with it. | Codex review on PR #1858; docs/filter-contract.md section 5 | 2026-08-12 |
| #310 | P2 | issue | Fuzzy catalogue search can match a DIFFERENT drug: fluoxetine to duloxetine at edit distance 2 | MEASURED 2026-08-12 by running the matcher itself, not by reading it. PR #1851 adds Damerau-Levenshtein typo recovery to `src/lib/catalog-search.ts` (`fuzzySearchTokenCount`, `boundedTypoDistance`, `typoDistanceLimit`) and folds it into the score. The tier `term.length >= 8 -> 2 edits` is the problem: Damerau counts an adjacent transposition as ONE edit, so `fluoxetine` -> `duloxetine` is distance 2 (substitute f->d, transpose lu->ul) and both are 10 characters. Confirmed hits against the PR's own algorithm: **fluoxetine -> duloxetine** (SSRI vs SNRI, different drugs), **prednisone -> prednisolone** (different drugs). Intended cases also confirmed working: sertraline -> setraline, olanzapine -> olanzepine. The existing guards DO hold — SSRI/SNRI, ADHD/ODD, citalopram/escitalopram, clozapine/clonazepam and quetiapine/olanzapine all correctly return no match. ONE MITIGATION, stated so this is not over-read: terms under 5 characters are excluded entirely. The fuzzy trigger is evaluated independently for each candidate record, so the hazard persists when both the exact drug and a two-edit near-match are present: the exact record receives a literal score while the wrong drug can independently receive a fuzzy score and appear as an additional result. Blast radius is wide because `catalog-search.ts` feeds ELEVEN modules — medications.ts (prescribing), dsm.ts, differentials.ts, differential-stream.ts, universal-search.ts, specifiers-search-index.ts, tools-catalog.ts, form-ranker.ts, service-ranker.ts. TESTED FIX: capping the >=8 tier at 1 edit removes both cross-drug hits and preserves every legitimate typo recovery in the sample — a one-line change to `typoDistanceLimit`. Next: if PR #1851 is still open, raise this on it; if it merged, apply the cap directly and add a test over real catalogue drug names with both the exact and near-match records present, asserting the wrong drug is excluded while the exact drug remains. Stop: do not remove fuzzy search outright — the typo recovery is genuinely useful and the guards are otherwise well judged. Note `classifyPullRequestFiles` returns clinicalRisk:true for this path (governance preflight fires) but ragRanking:false, which is correct — this is catalogue ranking, not the pgvector retrieval path. | session 2026-08-12; PR #1851 (codex/investigate-recent-regression-issues); algorithm re-run locally against real drug-name pairs; src/lib/catalog-search.ts | 2026-08-12 |
| #311 | P3 | task | Promote the derived ledger loss-detector into scripts/ — it has now earned its place twice | During the 2026-08-12 sweep, two main-merges silently reverted edits to `docs/outstanding-issues.md`, including the ENTIRE #293 refutation (a `grep sm:min-h-0` returned 0; the text survived only in commit a6bfc6f). It went unnoticed because the recovery script was HAND-ENUMERATED — it listed 15 archives and 8 updates from one commit and could therefore only restore what the author remembered. The replacement is derived rather than listed: read every row id this branch has ever stamped out of `git rev-list <merge-base>..HEAD` plus `git show <commit>:docs/outstanding-issues.md`, then assert each of those ids that is still OPEN carries its stamp text, and exit non-zero listing any that lost it. It has now proved itself twice — it caught the intentional #262 divergence (main's version was newer than the branch's, correctly left alone) and would have caught the #293 loss the hand-written list missed. The plan that created it said it should stay a scratch script 'unless it proves useful more than once'; that condition is met. Next: port it to scripts/ (suggested `check-ledger-stamp-retention.mjs`), generalise the stamp token from the hard-coded 2026-08-12 date to a `--since` or marker argument, add a self-test in the style of the other ledger scripts, and document it beside `ledger:dedupe` for use after any main sync that touches the ledger. Stop: do NOT wire it into verify:cheap or CI — it is a branch-local safety net for a human or agent mid-sweep, and it has no meaning on a branch that has not stamped rows. Related: #156 and #168, which track the id-allocation race that produces these merges in the first place. | session 2026-08-12 ledger sweep; scratch loss-check.mjs; #293 restoration from a6bfc6f | 2026-08-12 |

## Resolved / archive

Expand Down
Loading