Skip to content

docs(playbooks,claude-config): give the matrix rule a custody baseline and I8-c its second source - #1919

Merged
kyle-sexton merged 2 commits into
mainfrom
docs/roster-r13-thinking-troubleshooting
Aug 4, 2026
Merged

docs(playbooks,claude-config): give the matrix rule a custody baseline and I8-c its second source#1919
kyle-sexton merged 2 commits into
mainfrom
docs/roster-r13-thinking-troubleshooting

Conversation

@kyle-sexton

Copy link
Copy Markdown
Contributor

Summary

Doc-alignment roster row 13: Troubleshooting thinking — the per-model matrix page the calibration rule and the I17 family cite as authority, previously cited-but-not-digested (no hash, no capture). First byte-level baseline taken: 12,544 B, MD5 dc994aa9…, confirmed by producer and verifier on independent fetches.

playbooks 0.6.14calibration.md's matrix citation gains the custody its own re-check mandate requires: a capture stamp framed forward-only (the prior citation had no hash, so no unchanged-since claim is possible and none is made); the worked instance's two halves now date separately (matrix re-read 2026-08-04 and unmoved; introducing-page quote + registry reading remain 2026-08-03 snapshots); the verified negative tightened to name the page's only other availability sentence (the Claude 4 deprecations pointer) so it is falsifiable on its own terms.

claude-config 0.21.6 / criteria 1.14.0 — I8-c (don't-think directives increase tag leakage) gains its second source: this model-agnostic troubleshooting page states the same claim from the symptom side and still names only Claude Opus 5 — so the row's scope is now positively confirmed narrow on the declined-widening reasoning the catalog already uses (promotion gate deliberately unmet: the claim is not stated unqualified). The row also gains the consequence it lacked ("A leaked tool call never runs, and in agentic loops the leaked text stays in the conversation history…"), the condition ("most commonly on tool-heavy workloads such as search"), one catalog-extension sentence explicitly marked as the catalog's own, and a recheck trigger naming the two stating sections (an absence has no page to watch). Sources parenthetical updated.

All four pre-existing citations that rest on this page spot-verified against the live table (Mythos row, adaptive-only models, reject values, footnote 2's per-request enforcement) — nothing misstated. Deliberate non-actions with evidence: no reverse-direction-400 row (no instruction-text shape prescribes adaptive — it is the unspecified default), max_tokens truncation owned by the errors pages, display-defaults covered by I10 + context-economy. The page's docpage-digest queue position stands (queue rationale re-verified, strengthened by the page's own Next-steps transfer declaration).

Test plan

  • Docs-only; markdownlint 0 errors; changelog parity all three modes; instruction-scan.test.sh 46/46; conflict-scan.test.sh 41/41; scripted quote fidelity with control probes.
  • Orchestrator-commissioned Fable verifier: independent fetches of the troubleshooting page, the Opus 5 guide, and both errors pages (all byte-identical to producer captures); promotion-gate reading verified against the gate's own "stated unqualified" precedents; all three contested calls adjudicated sound; custody framing checked forward-only — 6/6 PASS, empty defect list.
  • Rebased onto post-feat(claude-config): flag a leading thinking block treated as required where none is #1918 main with a composed Sources block (both feat(claude-config): flag a leading thinking block treated as required where none is #1918's Steering entry and this row's expanded Troubleshooting entry preserved); stacks asserted: [0.21.6] > [0.21.5] > [0.21.4]; criteria 1.14.0.

Related

🤖 Generated with Claude Code

https://claude.ai/code/session_019gaVX25Txd6GXdiu9HEH3X

…and I8-c its second source

Roster row 13 — Troubleshooting thinking — was cited-but-not-digested: two shipped
surfaces rest on the page's per-model table and none of them had a capture behind the
citation. Live fetch 2026-08-04, HTTP 200, 12,544 B, MD5
dc994aa9129fbbebf0813f3241349971.

Every existing citation holds. calibration.md's authority claim matches the table's
actual column set; the Mythos 5 row exists; the criteria Sources parenthetical matches
the page's four 400 sections and footnote 2's model range; I17 base's quoted fragment
"is enforced on each request" is character-exact; I17-c's API-arm model list checks out
row by row, with Mythos Preview correctly excluded because it accepts extended thinking.
Nothing was misstated. What was missing was custody and one uncovered claim.

playbooks 0.6.14 — the matrix rule mandates a re-check trigger for any matrix a reader
acts on, and its own worked instance carried a trigger with nothing to re-check against.
The citation now carries the capture, and says plainly that it dates continuity forward
and claims none backward: this is the first byte-level capture of the page here, so no
earlier hash exists to compare with and no unchanged-since claim is available. The
instance's two halves now date separately — the matrix page re-read 2026-08-04 and its
Mythos 5 row unmoved, the introducing-page quote and the local registry reading still
2026-08-03 snapshots — rather than dating the whole instance to one day. Its verified
negative also names its scope: the page carries a second availability sentence (the
Claude 4 deprecations pointer), which does not weaken the negative but left an absolute
claim standing beside a literal counterexample.

claude-config 0.21.6 / criteria 1.14.0 — I8-c held Model scope: opus-5 only because no
wider statement had been found. This page states the same claim from the symptom side on
a model-agnostic feature page — the surface where a wider claim would appear — and names
Claude Opus 5 anyway, so the scope is now positively confirmed narrow on the same
declined-widening reasoning I10 already uses. The row also gains the consequence it
lacked: a leaked tool call never runs and the leaked text stays in agentic conversation
history, which makes this a history-poisoning failure in an autonomous lane rather than
one malformed response, plus the condition that tells an auditor where to look. Its
recheck trigger names the two sections that state the claim, because neither page
enumerates the models that do not leak — an absence has no page to watch, so the trigger
points at where a second model name would have to appear.

Deliberate non-actions: no row for the reverse-direction 400 (adaptive rejected on
extended-only models) — adaptive is the default wherever it is supported, so no
instruction text prescribes it, and the direction that does get prescribed is I17-c's;
the API-side reach belongs to the row-6 deferral's existing first-consumer trigger. No
row for the max_tokens truncation symptom (not an instruction-text shape; the harness
owns its own variant). display-defaults-to-omitted and effort-does-not-change-thinking
are already covered by I10 and context-economy. The existing I17 verification stamps
were left alone: they read 2026-08-04 and today's fetch makes them accurate as written.

VERSION NOTE: criteria 1.13.0 / claude-config 0.21.5 belong to open PR #1918. This
branch takes 1.14.0 / 0.21.6, skipping those. The changelog entry deliberately asserts
no predecessor version, since 1.13.0 is not on this branch and may never land.
Rebase before merge.

Verification: markdownlint 0 errors over all four changed files; changelog parity
--check, --check-order, --check-bump all pass with the 0.21.5 gap; instruction-scan
46/46 and conflict-scan 41/41; scripted quote fidelity over the I8-c row against the
capture — 5 quotes character-exact, 3 correctly attributed to the Opus 5 guide and
skipped, 3 altered control probes correctly absent; 7 scripted structural checks of the
calibration custody claims against the capture, including the falsifier that Glasswing
and limited-availability language are absent from this page, all pass.

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
@cursor

cursor Bot commented Aug 4, 2026

Copy link
Copy Markdown

Bugbot is not enabled for your account, so this pull request was not reviewed.

Enable Bugbot in the Cursor dashboard to get automatic reviews on future PRs.

@claude

claude Bot commented Aug 4, 2026

Copy link
Copy Markdown
Contributor

Claude finished @kyle-sexton's task in 45s —— View job


Security Review — PR #1919

  • Gather context (PR diff via git diff origin/main...HEAD)
  • Review changed files for security issues
  • Post findings

No security issues found.

This PR is documentation-only: two plugin.json version bumps, two CHANGELOG.md prose additions, and edits to two markdown reference/doctrine files (plugins/claude-config/skills/audit-instructions/reference/criteria.md, plugins/playbooks/skills/fable-5/context/calibration.md). There is no code, no shell scripts, no GitHub Actions workflow changes, no hook wiring, no .mcp.json / permissions / config changes, and no secrets or credentials in the diff.

Checked specifically for the classes called out in scope:

  • Injection surfaces (command/SQL/path/template): none — no executable code paths touched.
  • Tokens/secrets/credentials: none present; no MD5 dc994aa9… or other hash values are secrets, just a byte-capture fingerprint for doc-drift detection.
  • GitHub Actions risk (pull_request_target/workflow_run, script injection via github context, permission widening): N/A — no workflow files touched.
  • Supply-chain/unpinned actions: N/A — no action or dependency pins touched.
  • URLs added: all point to the existing official platform.claude.com docs domain already cited elsewhere in these files; nothing new or untrusted.

Nothing further to report for this change.

@claude

claude Bot commented Aug 4, 2026

Copy link
Copy Markdown
Contributor

Claude finished @kyle-sexton's task in 3m 8s —— View job


Review complete

Reviewed the diff (git diff origin/main...HEAD) across all 6 changed files: both plugin.json version bumps, both CHANGELOG.md entries, criteria.md (I8-c row), and calibration.md (the matrix-citation rule + Mythos 5 worked instance).

Mechanical checks: both plugin.json files are valid JSON with version bumps that match their changelogs (claude-config 0.21.5→0.21.6, playbooks 0.6.13→0.6.14, criteria frontmatter 1.13.0→1.14.0); .markdownlint-cli2.jsonc disables MD013 so the long prose lines aren't a lint issue; the new paragraph inserted directly under the Row I8-c header matches the existing Row I8-b precedent (line 297), and the **Verified DATE** / **Recheck trigger:** formatting matches the convention used throughout the rest of criteria.md (e.g. lines 448, 765, 845, 899, 1024).

One correctness finding:

Stale single-date stamp contradicts the split-dating this PR introducesplugins/playbooks/skills/fable-5/context/calibration.md#L65

The Mythos 5 worked instance opens with > Worked instance, verified 2026-08-03. — untouched by this PR. But the PR's own edit to the closing re-check-trigger line (L68) now says the matrix half was re-read 2026-08-04 while only the introducing-page quote and registry reading remain 2026-08-03 snapshots — explicitly "rather than dating the whole instance to one day" (per this PR's own changelog wording, plugins/playbooks/CHANGELOG.md L25). The opening stamp still does exactly that: it reads as a single verification date for the whole instance, which is now inaccurate for the matrix-derived content. A reader skimming just the header takes 2026-08-03 as authoritative for the entire block, missing the 2026-08-04 re-verification the PR added. Given this PR's explicit purpose is precise per-fact dating custody, the header should either be dropped (the per-half dating in L68 already carries the real information) or updated to reflect the split (e.g. "verified 2026-08-03, matrix re-read 2026-08-04").

Fix this →

Note: I did not independently re-fetch the vendor pages cited (the troubleshooting-thinking page's byte count/MD5, or its leakage-claim quote) — verifying those against live docs is outside a diff-scoped code review, and per CLAUDE.md's fresh-docs mandate carve-out this is prose-body content (not a manifest/schema/contract surface), so I relied on internal consistency rather than re-verification.

Everything else — the I8-c scope-narrowing logic, the new consequence/condition text, the Sources parenthetical update, and both changelog entries — is internally consistent and well-supported by the diff.

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: 1813801759

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

Comment thread plugins/claude-config/skills/audit-instructions/reference/criteria.md Outdated
@github-actions

github-actions Bot commented Aug 4, 2026

Copy link
Copy Markdown

Claude has reviewed this PR 2 times. The lane skips further automatic reviews after 5; deleting this comment resets the count.

Codex finding: the I8-c recheck trigger said a second model name "would
move the promotion gate", but a claim qualified to two models still
licenses nothing about the rest. The trigger now states the added model
joins as a named Detect condition and unscoping still requires what the
gate has always required - an unqualified model-agnostic statement or
convergent model guides.

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
@claude

claude Bot commented Aug 4, 2026

Copy link
Copy Markdown
Contributor

Claude finished @kyle-sexton's task in 2m 34s —— View job


Review complete

  • Gather context (diff vs origin/main, prior review comments)
  • Review changed files for correctness & alignment with CLAUDE.md/REVIEW.md
  • Post findings

Reviewed both commits (1813801 and the fix-up 3bb213b) across all 6 changed files. The codex P2 finding on the I8-c recheck trigger (a second named model shouldn't be read as satisfying the promotion gate) was correctly fixed in 3bb213bcriteria.md#L359-L365 now correctly says a new name "re-opens the scoping question, not the gate itself," and still requires "an unqualified model-agnostic statement, or convergent model guides" to unscope.

Two findings, one new:

  1. The 3bb213b fix wasn't propagated to the changelog it accompaniesplugins/claude-config/CHANGELOG.md#L28-L30

    The changelog entry still reads: "The Source line carries its own verification date and a recheck trigger keyed to the claim gaining a model beyond Opus 5, which is the event that would move the promotion gate." That's exactly the overstatement criteria.md was just corrected to not say — the fixed row text now explicitly states a second name does not move the gate ("not the gate itself... unscoping still requires what the gate has always required"). Since the changelog is the human-facing summary of what the row now claims, a reader who only skims the changelog comes away with the wrong rule. Only criteria.md was touched in the fix-up commit; the changelog prose that describes the pre-fix behavior was never updated to match.

    Fix this →

  2. Carried over from the prior review round, still unaddressedplugins/playbooks/skills/fable-5/context/calibration.md#L65

    The Mythos 5 worked instance still opens with the single-date stamp > Worked instance, verified 2026-08-03., while the closing re-check trigger at L68 now correctly splits dating: matrix re-read 2026-08-04, introducing-page quote and registry reading still 2026-08-03. The 3bb213b follow-up commit fixed the codex finding in criteria.md but didn't touch this file, so the header/body mismatch a previous review round flagged is still live. The playbooks changelog itself says the point is "rather than dating the whole instance to one day" — the header text still does exactly that to a reader who doesn't reach L68.

Everything else — the I8-c scope-narrowing logic, the new consequence/condition text, the Sources parenthetical update in both files, and the version/changelog bumps (claude-config 0.21.5→0.21.6, playbooks 0.6.13→0.6.14, criteria 1.13.0→1.14.0, all matching their changelogs) — is internally consistent and well-supported by the diff.

I did not independently re-fetch the vendor pages cited (byte counts, MD5, or quoted claim text) — that's outside a diff-scoped code review and, per CLAUDE.md's fresh-docs mandate carve-out, this is prose-body content rather than a manifest/schema/contract surface.

@claude

claude Bot commented Aug 4, 2026

Copy link
Copy Markdown
Contributor

Claude finished @kyle-sexton's task in 34s —— View job


Security Review — PR #1919

  • Gather context (git diff origin/main...HEAD)
  • Review changed files for security issues
  • Post findings

No security issues found.

This PR is documentation-only. All 6 changed files are: two plugin.json version bumps (0.21.5→0.21.6, 0.6.13→0.6.14), two CHANGELOG.md prose additions, and edits to two markdown reference/doctrine files (plugins/claude-config/skills/audit-instructions/reference/criteria.md, plugins/playbooks/skills/fable-5/context/calibration.md).

Checked specifically for the classes in scope:

  • Injection surfaces (command/SQL/path/template): none — no executable code, no shell scripts, no template rendering touched by this diff.
  • Tokens/secrets/credentials: none present. The MD5 dc994aa9… value is a content-integrity fingerprint of a public vendor doc page (used for drift detection), not a secret or credential.
  • GitHub Actions risk (pull_request_target/workflow_run with secrets over untrusted code, script injection via the github context, permission-widening): N/A — no workflow files (.github/workflows/**) are touched by this PR.
  • Supply-chain / unpinned actions: N/A — no action or dependency pins touched; both plugin.json diffs are version-string-only bumps consistent with their changelogs.
  • URLs added/changed: all point to the existing, already-cited platform.claude.com docs domain; nothing new or untrusted introduced, no data egress.
  • Config/permissions: no .mcp.json, settings.json, or permission-grant files touched.

No CRITICAL, IMPORTANT, or SUGGESTION-level security findings for this change.

@kyle-sexton
kyle-sexton merged commit 540f6a7 into main Aug 4, 2026
32 checks passed
@kyle-sexton
kyle-sexton deleted the docs/roster-r13-thinking-troubleshooting branch August 4, 2026 09:09
kyle-sexton added a commit that referenced this pull request Aug 4, 2026
…ive pages (#1920)

## Summary

Doc-alignment roster row 16: **Prompt caching (API)** — first captures
of both the API page (152,223 B) and its harness sibling (29,721 B),
with a hard surface-scoping discipline (API vs harness caching semantics
are different products' claims). Seven repo surfaces stating caching
facts swept; five verified clean and left alone (including boris's
$12.50/$1 figures — confirmed contextually right for subagent
orchestration, where the five-minute TTL governs); two carried real
defects:

**playbooks 0.6.15** — `orchestration.md`'s continue-an-oriented-worker
rationale claimed "accumulated context is a cache read". Wrong in the
chapter's own modal case: the harness page states subagents build their
own cache *and* use the five-minute TTL even on subscription, so a
worker resumed after a longer fan-out wave re-writes its whole context
at the five-minute cache-write rate ("1.25 times the base input tokens
price"), not a cache-read rate. The recommendation stands; the reason is
now the re-derivation saved (a replacement pays the same tokens plus the
rediscovery tool turns), with the TTL and pricing anchors cited.

**docs-hygiene 0.9.5** — `extract-ssot`'s anti-pattern #9 ("Cache
invalidation cascade") rested on a mechanism that is dead on the skill's
own declared surface: the harness page states mid-session edits of
always-loaded files keep the cache (the edit just doesn't apply), and
cross-session sharing keys on the git-status snapshot, which any commit
breaks. Rewritten in place as **"Always-loaded SSOT propagation lag"** —
corrections ship that live sessions don't see until
`/clear`/`/compact`/restart — with a scope fence for the API surface
(where prefix volatility genuinely costs an Agent SDK fleet), scoped to
*unscoped* rules files (path-scoped rules load lazily; pre-load edits
apply), and slot 9 preserved because #10#13 are cited by number in
eight places. The dead vocabulary survives only in the changelog, quoted
as removed.

Routed, not acted on: the owner's dotfiles CLAUDE.md caching claim
verified correct with one additive omission (fast mode is a third
cache-key element) — recorded for dotfiles routing; the API page's
explicit-default-effort no-invalidate row flagged as a future I17-b
enrichment parallel.

## Test plan

- Docs-only; markdownlint 0 errors; version/changelog parity both
plugins; scripted quote fidelity 11/11 against the captured pages.
- Orchestrator-commissioned Fable verifier: both live fetches (MD5s
exact), both correction logics reconstructed, the slot-9 citation count
independently verified, both contested calls upheld — substance PASS
with 1 real defect (a surviving dead-mechanism table row in the same
file) + 2 precision nits (cache-write rate; unscoped-rules scoping), all
three fixed and re-checked **ALL PASS** including an independent
dead-vocabulary sweep.

## Related

- No linked issue.
- Doc-alignment loop, roster row 16. Predecessors: #1908#1919.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

https://claude.ai/code/session_019gaVX25Txd6GXdiu9HEH3X

---------

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant