docs(playbooks,claude-config): give the matrix rule a custody baseline and I8-c its second source - #1919
Conversation
…and I8-c its second source Roster row 13 — Troubleshooting thinking — was cited-but-not-digested: two shipped surfaces rest on the page's per-model table and none of them had a capture behind the citation. Live fetch 2026-08-04, HTTP 200, 12,544 B, MD5 dc994aa9129fbbebf0813f3241349971. Every existing citation holds. calibration.md's authority claim matches the table's actual column set; the Mythos 5 row exists; the criteria Sources parenthetical matches the page's four 400 sections and footnote 2's model range; I17 base's quoted fragment "is enforced on each request" is character-exact; I17-c's API-arm model list checks out row by row, with Mythos Preview correctly excluded because it accepts extended thinking. Nothing was misstated. What was missing was custody and one uncovered claim. playbooks 0.6.14 — the matrix rule mandates a re-check trigger for any matrix a reader acts on, and its own worked instance carried a trigger with nothing to re-check against. The citation now carries the capture, and says plainly that it dates continuity forward and claims none backward: this is the first byte-level capture of the page here, so no earlier hash exists to compare with and no unchanged-since claim is available. The instance's two halves now date separately — the matrix page re-read 2026-08-04 and its Mythos 5 row unmoved, the introducing-page quote and the local registry reading still 2026-08-03 snapshots — rather than dating the whole instance to one day. Its verified negative also names its scope: the page carries a second availability sentence (the Claude 4 deprecations pointer), which does not weaken the negative but left an absolute claim standing beside a literal counterexample. claude-config 0.21.6 / criteria 1.14.0 — I8-c held Model scope: opus-5 only because no wider statement had been found. This page states the same claim from the symptom side on a model-agnostic feature page — the surface where a wider claim would appear — and names Claude Opus 5 anyway, so the scope is now positively confirmed narrow on the same declined-widening reasoning I10 already uses. The row also gains the consequence it lacked: a leaked tool call never runs and the leaked text stays in agentic conversation history, which makes this a history-poisoning failure in an autonomous lane rather than one malformed response, plus the condition that tells an auditor where to look. Its recheck trigger names the two sections that state the claim, because neither page enumerates the models that do not leak — an absence has no page to watch, so the trigger points at where a second model name would have to appear. Deliberate non-actions: no row for the reverse-direction 400 (adaptive rejected on extended-only models) — adaptive is the default wherever it is supported, so no instruction text prescribes it, and the direction that does get prescribed is I17-c's; the API-side reach belongs to the row-6 deferral's existing first-consumer trigger. No row for the max_tokens truncation symptom (not an instruction-text shape; the harness owns its own variant). display-defaults-to-omitted and effort-does-not-change-thinking are already covered by I10 and context-economy. The existing I17 verification stamps were left alone: they read 2026-08-04 and today's fetch makes them accurate as written. VERSION NOTE: criteria 1.13.0 / claude-config 0.21.5 belong to open PR #1918. This branch takes 1.14.0 / 0.21.6, skipping those. The changelog entry deliberately asserts no predecessor version, since 1.13.0 is not on this branch and may never land. Rebase before merge. Verification: markdownlint 0 errors over all four changed files; changelog parity --check, --check-order, --check-bump all pass with the 0.21.5 gap; instruction-scan 46/46 and conflict-scan 41/41; scripted quote fidelity over the I8-c row against the capture — 5 quotes character-exact, 3 correctly attributed to the Opus 5 guide and skipped, 3 altered control probes correctly absent; 7 scripted structural checks of the calibration custody claims against the capture, including the falsifier that Glasswing and limited-availability language are absent from this page, all pass. Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
|
Bugbot is not enabled for your account, so this pull request was not reviewed. Enable Bugbot in the Cursor dashboard to get automatic reviews on future PRs. |
|
Claude finished @kyle-sexton's task in 45s —— View job Security Review — PR #1919
No security issues found. This PR is documentation-only: two Checked specifically for the classes called out in scope:
Nothing further to report for this change. |
|
Claude finished @kyle-sexton's task in 3m 8s —— View job Review complete
Reviewed the diff ( Mechanical checks: both One correctness finding: Stale single-date stamp contradicts the split-dating this PR introduces — The Mythos 5 worked instance opens with Note: I did not independently re-fetch the vendor pages cited (the troubleshooting-thinking page's byte count/MD5, or its leakage-claim quote) — verifying those against live docs is outside a diff-scoped code review, and per CLAUDE.md's fresh-docs mandate carve-out this is prose-body content (not a manifest/schema/contract surface), so I relied on internal consistency rather than re-verification. Everything else — the I8-c scope-narrowing logic, the new consequence/condition text, the |
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: 1813801759
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
|
Claude has reviewed this PR 2 times. The lane skips further automatic reviews after 5; deleting this comment resets the count. |
Codex finding: the I8-c recheck trigger said a second model name "would move the promotion gate", but a claim qualified to two models still licenses nothing about the rest. The trigger now states the added model joins as a named Detect condition and unscoping still requires what the gate has always required - an unqualified model-agnostic statement or convergent model guides. Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
|
Claude finished @kyle-sexton's task in 2m 34s —— View job Review complete
Reviewed both commits ( Two findings, one new:
Everything else — the I8-c scope-narrowing logic, the new consequence/condition text, the I did not independently re-fetch the vendor pages cited (byte counts, MD5, or quoted claim text) — that's outside a diff-scoped code review and, per CLAUDE.md's fresh-docs mandate carve-out, this is prose-body content rather than a manifest/schema/contract surface. |
|
Claude finished @kyle-sexton's task in 34s —— View job Security Review — PR #1919
No security issues found. This PR is documentation-only. All 6 changed files are: two Checked specifically for the classes in scope:
No CRITICAL, IMPORTANT, or SUGGESTION-level security findings for this change. |
…ive pages (#1920) ## Summary Doc-alignment roster row 16: **Prompt caching (API)** — first captures of both the API page (152,223 B) and its harness sibling (29,721 B), with a hard surface-scoping discipline (API vs harness caching semantics are different products' claims). Seven repo surfaces stating caching facts swept; five verified clean and left alone (including boris's $12.50/$1 figures — confirmed contextually right for subagent orchestration, where the five-minute TTL governs); two carried real defects: **playbooks 0.6.15** — `orchestration.md`'s continue-an-oriented-worker rationale claimed "accumulated context is a cache read". Wrong in the chapter's own modal case: the harness page states subagents build their own cache *and* use the five-minute TTL even on subscription, so a worker resumed after a longer fan-out wave re-writes its whole context at the five-minute cache-write rate ("1.25 times the base input tokens price"), not a cache-read rate. The recommendation stands; the reason is now the re-derivation saved (a replacement pays the same tokens plus the rediscovery tool turns), with the TTL and pricing anchors cited. **docs-hygiene 0.9.5** — `extract-ssot`'s anti-pattern #9 ("Cache invalidation cascade") rested on a mechanism that is dead on the skill's own declared surface: the harness page states mid-session edits of always-loaded files keep the cache (the edit just doesn't apply), and cross-session sharing keys on the git-status snapshot, which any commit breaks. Rewritten in place as **"Always-loaded SSOT propagation lag"** — corrections ship that live sessions don't see until `/clear`/`/compact`/restart — with a scope fence for the API surface (where prefix volatility genuinely costs an Agent SDK fleet), scoped to *unscoped* rules files (path-scoped rules load lazily; pre-load edits apply), and slot 9 preserved because #10–#13 are cited by number in eight places. The dead vocabulary survives only in the changelog, quoted as removed. Routed, not acted on: the owner's dotfiles CLAUDE.md caching claim verified correct with one additive omission (fast mode is a third cache-key element) — recorded for dotfiles routing; the API page's explicit-default-effort no-invalidate row flagged as a future I17-b enrichment parallel. ## Test plan - Docs-only; markdownlint 0 errors; version/changelog parity both plugins; scripted quote fidelity 11/11 against the captured pages. - Orchestrator-commissioned Fable verifier: both live fetches (MD5s exact), both correction logics reconstructed, the slot-9 citation count independently verified, both contested calls upheld — substance PASS with 1 real defect (a surviving dead-mechanism table row in the same file) + 2 precision nits (cache-write rate; unscoped-rules scoping), all three fixed and re-checked **ALL PASS** including an independent dead-vocabulary sweep. ## Related - No linked issue. - Doc-alignment loop, roster row 16. Predecessors: #1908–#1919. 🤖 Generated with [Claude Code](https://claude.com/claude-code) https://claude.ai/code/session_019gaVX25Txd6GXdiu9HEH3X --------- Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
Summary
Doc-alignment roster row 13: Troubleshooting thinking — the per-model matrix page the calibration rule and the I17 family cite as authority, previously cited-but-not-digested (no hash, no capture). First byte-level baseline taken: 12,544 B, MD5
dc994aa9…, confirmed by producer and verifier on independent fetches.playbooks 0.6.14 —
calibration.md's matrix citation gains the custody its own re-check mandate requires: a capture stamp framed forward-only (the prior citation had no hash, so no unchanged-since claim is possible and none is made); the worked instance's two halves now date separately (matrix re-read 2026-08-04 and unmoved; introducing-page quote + registry reading remain 2026-08-03 snapshots); the verified negative tightened to name the page's only other availability sentence (the Claude 4 deprecations pointer) so it is falsifiable on its own terms.claude-config 0.21.6 / criteria 1.14.0 — I8-c (don't-think directives increase tag leakage) gains its second source: this model-agnostic troubleshooting page states the same claim from the symptom side and still names only Claude Opus 5 — so the row's scope is now positively confirmed narrow on the declined-widening reasoning the catalog already uses (promotion gate deliberately unmet: the claim is not stated unqualified). The row also gains the consequence it lacked ("A leaked tool call never runs, and in agentic loops the leaked text stays in the conversation history…"), the condition ("most commonly on tool-heavy workloads such as search"), one catalog-extension sentence explicitly marked as the catalog's own, and a recheck trigger naming the two stating sections (an absence has no page to watch). Sources parenthetical updated.
All four pre-existing citations that rest on this page spot-verified against the live table (Mythos row, adaptive-only models, reject values, footnote 2's per-request enforcement) — nothing misstated. Deliberate non-actions with evidence: no reverse-direction-400 row (no instruction-text shape prescribes adaptive — it is the unspecified default), max_tokens truncation owned by the errors pages, display-defaults covered by I10 + context-economy. The page's docpage-digest queue position stands (queue rationale re-verified, strengthened by the page's own Next-steps transfer declaration).
Test plan
instruction-scan.test.sh46/46;conflict-scan.test.sh41/41; scripted quote fidelity with control probes.Related
api/errors, 22,669 B; harnesserrors, 152,674 B), both now captured.🤖 Generated with Claude Code
https://claude.ai/code/session_019gaVX25Txd6GXdiu9HEH3X