Skip to content

Harden RAG telemetry and production guardrails - #136

Merged
BigSimmo merged 6 commits into
mainfrom
bigsimmo-ai-architecture-review
Jul 2, 2026
Merged

Harden RAG telemetry and production guardrails#136
BigSimmo merged 6 commits into
mainfrom
bigsimmo-ai-architecture-review

Conversation

@BigSimmo

@BigSimmo BigSimmo commented Jul 2, 2026

Copy link
Copy Markdown
Owner

Summary

  • add explicit degraded-mode signaling across answer/search API payloads
  • expand shared-cache observability with miss reason diagnostics
  • add prompt-injection false-positive eval coverage and route-level p95 latency thresholds
  • add production readiness fail-fast when raw query persistence is enabled

Validation

  • npm test

- Add explicit degraded-mode signaling across answer/search APIs\n- Expand RAG telemetry with shared-cache miss diagnostics\n- Add prompt-injection eval cases and route-level latency thresholds\n- Fail production readiness when raw query text persistence is enabled\n\nCo-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: 958ed42def

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

Comment thread src/lib/rag-eval-cases.ts Outdated
Comment thread src/lib/rag.ts
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: 091eef45d9

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

Comment thread src/app/api/answer/route.ts Outdated
BigSimmo and others added 3 commits July 2, 2026 17:02
The clozapine injection control deterministically classifies as medication_dose_risk, so align expectedQueryClass to avoid false query-class mismatches in eval quality gates.\n\nCo-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: 86f722f0ee

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

Comment thread src/lib/rag.ts
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
@BigSimmo
BigSimmo enabled auto-merge July 2, 2026 10:20
@BigSimmo
BigSimmo merged commit 59e4a93 into main Jul 2, 2026
4 checks passed

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: 9324c66a88

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

Comment thread src/lib/rag.ts
Comment on lines +7245 to +7246
shared_cache_status: search.telemetry.shared_cache_status,
shared_cache_miss_reason: search.telemetry.shared_cache_miss_reason,

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Badge Clear shared-cache miss fields from answer cache hits

When this generated answer is stored by setCachedAnswer below, a first request that missed the shared search cache persists shared_cache_status: "miss" and the miss reason inside answer.latencyTimings. Later exact-query answer-cache hits go through getCachedAnswer, which only refreshes total_latency_ms, so those requests are still reported as shared-cache misses even though no shared-cache lookup happened; clear these fields before answer-cache storage or when serving answer-cache hits.

Useful? React with 👍 / 👎.

Comment thread src/lib/rag.ts
Comment on lines 5207 to +5209
const answerQualityTier: RagAnswer["answerQualityTier"] =
answer.answerQualityTier ??
(answer.modelUsed ? "model_synthesis" : answer.routingMode === "extractive" ? "source_only" : undefined);
(answer.modelUsed ? "model_synthesis" : inferredSourceOnlyFallback ? "source_only" : undefined);

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Badge Mark extractive recovery answers as degraded

When a fast generation is discarded and recovered with buildExtractiveAnswer, the fallback keeps routingMode: "extractive" but line 7310 later restores modelUsed to the attempted model. In that scenario this ternary labels the response as model_synthesis, so the new degradedMode is inactive even though the client received a source-only extractive fallback; check the extractive/source-only route before modelUsed, or set answerQualityTier on the fallback explicitly.

Useful? React with 👍 / 👎.

@BigSimmo
BigSimmo deleted the bigsimmo-ai-architecture-review branch July 2, 2026 16:31
BigSimmo pushed a commit that referenced this pull request Jul 30, 2026
…ocation

CI caught what I did not: `static-pr` failed on `check:outstanding-issues` with
#131-#134 duplicated and two `issues:next-id` markers.

Cause: main's PR #1424 allocated #131-#134 for its own findings at the same time
this branch held #131-#135, and `merge=union` did what union does — kept both
sides under the same ids. That is #112's documented limit: union preserves
concurrent appends but cannot allocate unique ids, so the structural gate is the
only thing that catches it.

My error was pushing without re-running that gate. The previous push resolved a
`docs/branch-review-ledger.md` conflict, and I validated only that file before
pushing to win the race against main — but the same merge also touched
`docs/outstanding-issues.md`. `verify:cheap` would have caught it locally.

Main's rows keep #131-#134 (already merged and referenced elsewhere); this
branch's five renumber to #136-#140, one marker at 141, and the cold-cache
cross-reference in process-hardening follows its row.

Two of main's new rows also make a planned addition here redundant: #134 is the
absent ledger merge driver and #133 is the outstanding-issues merge churn — both
hit during this branch's work, both already captured upstream, so nothing new is
filed for them.
BigSimmo added a commit that referenced this pull request Jul 30, 2026
…t did (#1427)

* test(phone-scroll): prove the drag delivered before asserting the chrome hid

CI run 30518866604 failed `ui-phone-scroll.spec.ts:423` on

  expect(getByTestId('universal-header-collapse'))
    .toHaveAttribute('data-scroll-hidden', 'true')  // received ""

after the full 10s auto-retry, and the classifier recorded it as "needs
investigation". The assertion was right; the scroll never happened.

`dragScrollBy` moved the scroller with `scrollTop +=`, which clamps silently
at the end of the range, and returned nothing. When a page lays out shorter
than the test assumed — content still settling under full-suite CI load — a
720px request delivers a fraction of that, the chrome correctly stays visible
because document-detail chrome only hides past `scrollTop > 120`, and the
failure surfaces ten seconds later looking like a product regression. The
helper also resolved the scroll owner once up front, so a mid-drag layout
change left it pushing an element that had stopped scrolling.

- `dragScrollBy` now re-resolves the owner each step and returns the distance
  actually travelled.
- `dragScrollUntilHidden` waits for the remaining downward runway (a condition
  wait, not a settle sleep), drags, and fails naming the shortfall if the drag
  could not cross the threshold. Used at the four sites that assert a hide
  immediately after a fixed-distance drag.
- `addPhoneScrollRunway` waits for its 1600px filler to reach layout instead of
  sleeping 50ms. All 14 call sites already depend on that runway existing.

Every assertion is byte-identical: a genuinely stuck header still fails exactly
as before, once the drag is proven to have happened. No `.first()` was added
(#93's stop rule) and no tolerance was relaxed.

* ci: shard Production UI across three runners

Measured on 2026-07-30 from the Actions API, two full UI-scope PR runs
(30520443076, 30519912667): `Production UI` took 15m26-16m31 of a 16.8-18.6
minute run — 83-89% of wall clock — while every other job finished by minute 4
and then waited. Playwright itself reported `339 passed (13.5m)`; the balance is
the isolated production build.

That single job is also where the churn cost lands: 42% of PR runs in the
sampled window were cancelled (25 of 60 completed), almost all superseded
mid-Production-UI.

Sharding is across runners, not workers. `workers: 1`, `fullyParallel: false`
and `retries: 0` are unchanged inside each shard, so determinism is identical
and per-runner load falls — which matters because #93's duplicate page root is
load-dependent. `run-playwright.mjs` already forwards argv to `playwright test`,
so `--shard` needed no runner change.

The shard count is measured, not chosen. `fullyParallel: false` makes a spec
file indivisible, so shard sizes are lumpy and more shards is not monotonically
faster. Over the 340 required chromium tests:

  N=3 -> 121/106/113       largest 121
  N=4 -> 121/106/96/17     largest 121  (same critical path, one more runner)
  N=6 -> 65/56/106/5/91/17 largest 106
  N=5 -> 121/106/0/96/17   and N=8 -> two empty shards

N=4 buys nothing over N=3, and any N with an empty shard would go red because
`test:e2e:pr` deliberately omits `--pass-with-no-tests`. Expected critical path
~15.5 -> ~7 min, assuming per-test cost is roughly uniform.

`fail-fast: false` so a failing shard cannot cancel its siblings and re-create
the cancelled-vs-failed ambiguity #95 removed. Artifact names are shard-scoped
because upload-artifact runs with `overwrite: false`. Branch protection requires
only the `pr-required` aggregate, and `needs` on a matrix job yields the roll-up
of all shards, so the aggregate is unchanged.

Also adds `restore-keys` to both Playwright browser caches: without a prefix
fallback a lockfile bump forced a cold browser download in every UI job at once,
now three times over.

* ci: bound the codex auto-resolve jobs and serialise the visual config

Two inconsistencies found while mapping the pipeline, neither load-bearing but
both silent:

- `codex-autofix-review-comments.yml` was the only workflow in the repo with no
  `timeout-minutes` on either job, so both inherited GitHub's 360-minute default
  for work that reads PR metadata and posts one comment.
- `playwright.visual.config.ts` set neither `workers` nor `fullyParallel`, so it
  inherited Playwright's default `workers = 50% of CPUs`. The production config
  pins both to serial deliberately; the visual lane was quietly opting out of
  the anti-flake posture the rest of the suite is configured for.

* chore(gates): pin the documented gate count to the real chain

Both numbers were wrong. `CLAUDE.md` said 24 static/consistency gates against an
actual 25 — `check:assets` landed before that line was written, so it was wrong
at authoring — and the `gates` skill said "check 2 of 26" against an actual 28.

A stale count is not cosmetic here. The skill's whole point at that line is that
`verify:cheap` stops at the first failure and everything after it never ran; an
agent that believes the chain is 26 long cannot say how much a mid-chain failure
skipped.

`check:gate-manifest` already derives the real count from
`verify:cheap:internal`, so it now asserts the documented numbers against it.
The assertions fail closed: if the anchor phrasing disappears, the guard reports
a lost anchor rather than passing on a document it no longer checks.

Mutation-proven: reverting the skill to "26" fails with
".claude/skills/gates/SKILL.md says 26 where the chain has 28".

* docs(issues): capture the CI review's deferred findings

Five items from the CI/testing review that should not be changed blind:

- #125 `ui_changed` matches all of `src/app`, so an API-only diff pays the
  15-minute UI gate. Narrowing it can hide a real regression, so it needs a
  decision plus a compensating check rather than a quieter filter.
- #126 the Playwright build writes to a per-run distDir, so Next's build cache
  is cold every run (~2 min, now ~29% of the sharded critical path). Fixing it
  means suppressing the runner's documented always-cleanup, which must not ship
  without executing the runner.
- #127 the advisory UI lane spends ~3 min per UI PR on 5 mockup tests; there are
  currently zero `@quarantine` tests for it to cover.
- #128 CI Triage is complete and self-tested but inert pending a repo variable.
- #129 four `changes` outputs are computed and consumed by nothing, and
  `coverage_changed` fires on any non-doc file.

* docs(ledger): record the ci-testing-review pass at this HEAD

* ci: re-measure the shard split on the merged tree and refresh stale gate counts

The merge changed both numbers this branch had recorded.

Shard balance, re-measured against 342 required chromium tests (was 340):
  N=3 -> 121/111/110    largest 121
  N=4 -> 121/106/98/17  largest 121
N=3 remains correct — one 121-test spec group bounds both, so N=4 spends an
extra runner for the same critical path. The re-measure command is now in the
workflow comment so the next person does not have to rediscover it.

Gate counts: merging main added `check:gitleaks-pinned` and
`check:pr-mergeability` to `verify:cheap:internal`, so the documented counts
went stale the moment the merge landed — 25 -> 27 static, 28 -> 30 total. The
guard added earlier in this branch caught it immediately rather than letting the
docs drift again, which is the whole reason it exists.

Also records the `ui-critical-fast` interaction: the UI critical path is now that
15-test fail-fast job plus the slowest shard, not the full 13.5-minute suite, so
neither of this branch's pre-merge timings can be read on its own.

* docs(issues): rebuild the ledger after a union-merge duplication

The `merge=union` driver on `docs/outstanding-issues.md` preserves concurrent
appends, but when both sides restructure the same region it concatenates them
wholesale. Merging the latest main did exactly that: every open row appeared
twice and both `issues:next-id` markers survived — 66 duplicate-id errors from
`check:outstanding-issues`, which is precisely the failure that gate exists to
catch (#112).

Resolved by rebuilding on main's canonical file rather than by hand-editing the
duplicated table: reset to `origin/main`, then re-apply this branch's five
captured rows at #131-#135 (main had advanced its allocation to #130 while this
branch was open, so the earlier #128-#132 numbering collided again) and
re-apply the #127 narrowing note. Marker bumped to 136.

Union merge cannot allocate unique ids; only the structural gate can catch when
it has produced an invalid file. It did.

* ci: record the measured shard result, correcting the predicted one

First real run of the sharded shape (CI 30530618838, all green, whole run
13m39 against a 16.8-18.6 min unsharded baseline):

  ui-critical-fast  15 tests   3m14
  Production UI (1) 121 tests  9m36
  Production UI (2) 111 tests  6m54
  Production UI (3) 110 tests  6m20

The prediction was wrong by ~40%. ~6.8 min was expected for the largest shard
from 121/342 tests x 13.5 min; 9m36 happened. Per-test cost is not uniform —
111 tests took 6m54 while 121 took 9m36 — so a count-balanced split understates
the slowest shard whenever the slow specs land in one group. `--shard` can only
balance by count; balancing by duration would mean splitting the slow spec files
themselves.

The win is real but smaller than claimed, and the workflow comment and
process-hardening now carry the measured numbers plus the reason the arithmetic
misleads, so the next person re-measures instead of re-deriving.

Also merges origin/main. The ledger conflict was GitHub-visible only: that file
carries merge=union locally, which GitHub does not honour (#129). Resolved by
keeping the one genuinely new record and dropping three that main already had
elsewhere in the file — append-only forbids dropping a record that exists once,
not keeping a second copy. Superseding record appended for this HEAD, since the
prior one asserted a root cause that #127's trace evidence refutes.

* docs(issues): renumber this branch's rows above main's concurrent allocation

CI caught what I did not: `static-pr` failed on `check:outstanding-issues` with
#131-#134 duplicated and two `issues:next-id` markers.

Cause: main's PR #1424 allocated #131-#134 for its own findings at the same time
this branch held #131-#135, and `merge=union` did what union does — kept both
sides under the same ids. That is #112's documented limit: union preserves
concurrent appends but cannot allocate unique ids, so the structural gate is the
only thing that catches it.

My error was pushing without re-running that gate. The previous push resolved a
`docs/branch-review-ledger.md` conflict, and I validated only that file before
pushing to win the race against main — but the same merge also touched
`docs/outstanding-issues.md`. `verify:cheap` would have caught it locally.

Main's rows keep #131-#134 (already merged and referenced elsewhere); this
branch's five renumber to #136-#140, one marker at 141, and the cold-cache
cross-reference in process-hardening follows its row.

Two of main's new rows also make a planned addition here redundant: #134 is the
absent ledger merge driver and #133 is the outstanding-issues merge churn — both
hit during this branch's work, both already captured upstream, so nothing new is
filed for them.

* docs(issues): rebuild against main's current id allocation

The union merge duplicated the whole open and archive tables again (two header
rows, every id twice) because main restructured the file while this branch held
rows in it. Same resolution as before and for the same reason: rebuild on main's
canonical file rather than hand-editing a doubled table, then re-apply this
branch's five rows.

Main is now at next-id=135, so they land as #135-#139 with the marker at 140.
None of the five is duplicated upstream — checked by summary before re-applying.

This is the third renumber of the same five rows in one PR. That is not a
mistake being repeated, it is #133 ("outstanding-issues conflicts on nearly
every main advance") happening: any branch that holds rows in this file
re-collides every time main lands one. Worth weighing whether captures should
land in their own PR ahead of the work rather than riding along with it.

* docs: record PR 1427 review

---------

Co-authored-by: Claude <noreply@anthropic.com>
BigSimmo added a commit that referenced this pull request Jul 30, 2026
* docs: close rejected Playwright cache proposal

* docs: record PR 1467 review

* docs: record PR 1467 current-main review
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant