Skip to content

feat: apply the Opus 5.5 usage guide across repo instructions and plugins - #4352

Merged
kyle-sexton merged 10 commits into
mainfrom
feat/apply-opus-5-5-guide
Sep 23, 2026
Merged

kyle-sexton merged 10 commits into
mainfrom
feat/apply-opus-5-5-guide

Conversation

@kyle-sexton

Copy link
Copy Markdown
Contributor

No related issue: initiated directly from a session request to apply the Opus 5.5 usage guide across the repository; follow-ups are filed as #4346, #4347, #4348, #4349, and #4351.

Summary

Applies "Getting the most out of Opus 5.5 in Claude and Claude Code" (claude.dev, 2026-09-22) to this repository's own instructions and to every plugin's skills, agents, hook text, and prompts. The branch also adds an Opus 5.5 adaptation chapter to playbooks. docs/upstream/opus-5-5-usage-guide.md records each guide item with its verdict and the files that carry it.

Where the guide's advice is Opus-5.5-specific, it lives only in plugins/playbooks/reference/model-adaptation/opus-5-5.md. Shared skills carry only the model-agnostic items, because Fable 5.1, Sonnet, and Haiku also read them.

Fix

  • Repo instructions: the root AGENTS.md now says when to keep going and when to stop and ask. It asks the agent to stop before destructive actions or anything outside this checkout, keep permission prompts on, keep a task file on long runs, and end with "Blocked on me / Changed / Found" unless a skill defines its own report shape. The loop-lane launch prompts forbid ending a turn on a status summary.
  • Instruction audits:
    • claude-config:audit-instructions detects:
      • think-carefully steers, in the new I8-f row, scoped to Opus 5.5 targets;
      • requests to reproduce reasoning, in the widened I10 row;
      • vague design steers, in the extended I26 row;
      • misplaced settled-answer instructions, in the new I35 row.
    • audit-prompting-postures P1, P6, P9, and the new P11 check for finish lines, named stops, task files, fan-out evidence checks, and report order.
    • docs-hygiene:write-for-agents teaches the same items at authoring time.
  • playbooks: there is a new opus-5-5.md chapter, and fable-5 meta-rule 3 now routes it and re-resolves the chapter after a flagged-message fallback. opus-5.md and opus-4-8.md stay, because flagged Fable and Opus 5.5 requests fall back to Opus 5 (biology) and Opus 4.8 (cybersecurity). Each of those chapters now says so at its top, and playbooks: retire the Opus 5 and Opus 4.8 chapters once they stop being fallback targets #4349 tracks retiring them.
  • Long-run and orchestration:
    • implement-dispatch, lane-stop-gate, and the session-flow skills no longer treat a status summary or an offer to continue as a stop. No confirmation gate was weakened.
    • Spawn specs and dispatch briefs name what done means and when to stop early.
    • Fan-outs check each worker's evidence before accepting it.
  • Review and research:
    • Findings carry file and line, why the code is wrong, and how to show it fails.
    • PR prep and quality-gate PR mode lead with merge-blocking findings.
    • Audit reports lead with what waits on the user.
    • doc-drift-detector checks documents for contradictions within themselves.
    • Research marks what it could not confirm.
  • Design: visualization, prototype, playgrounds, education, and adhd name specific styles to leave out and extend that list when the user dislikes a choice.
  • Model currency: the live docs now say opus resolves to Opus 5.5, and known-issues covers the flagged-message model switch.

Every changed plugin is version-bumped and has a new CHANGELOG entry.

Decisions made by the owner:

  • review:code-review keeps its "block or flag" bar.
  • The visualization chrome keeps its look.
  • I8-f stays scoped to Opus 5.5 targets.

Verification

  • scripts/validate-plugins.sh: passes, including the strict catalog manifest.

  • scripts/check-changelog-parity.sh --check-bump origin/main and --check-preserved origin/main both pass: 37 changelogs were changed and 2612 headings were compared.

  • instruction-scan.test.sh passes 97/97 and emit-findings.test.sh passes 119/119.

  • lane-stop-gate.test.sh passes 103 of 104. The one failure is case 52, CR handling, which fails the same way at base on Windows.

  • check-skill.sh passes on the changed skills except for failures that already exist on main:

    • The descriptions of mutation-testing:audit, architecture:improve, and code-tidying:audit-dead-code are over the length limit (Three skill descriptions exceed the 1024-codepoint limit and fail check-skill #4351). The branch changes none of them.
    • Script tests fail in claude-config audit-automation-gaps (findings-state.test.sh, inventory.test.sh), claude-memory audit (instruction-load-stats.test.sh, nested-agents-check.test.sh), claude-ops changelog (changelog-status.test.sh), and repo-hygiene clean (git-branch-audit.test.sh). None of these tests read the files this branch edits. findings-state.test.sh and nested-agents-check.test.sh were confirmed to exit 1 on main too.
  • A fresh-context verifier checked the ledger against the diff. It found five problems:

    • PR prep dropped non-blocking findings.
    • implement-dispatch had two stop rules at a phase boundary.
    • The ledger attributed some G2.1 and G3.1 surfaces wrongly.
    • The guide's summary items had no mapping in the ledger.
    • Two wording conflicts between files.

    All five are fixed in 4a0c6e7dc. The verifier found no weakened gate, no new request to show reasoning, no broken ledger link, and no unbumped plugin.

  • markdownlint was not run locally because there is no node_modules here. CI runs it.

Related

🤖 Generated with Claude Code

https://claude.ai/code/session_015zXN1hq49epZqs3APTdCz8

kyle-sexton and others added 9 commits September 23, 2026 08:09
Adds reference/model-adaptation/opus-5-5.md sourced from the Opus 5.5
usage guide and the official prompting and model-config docs, routes it
from fable-5 meta-rule 3, and keeps the Opus 5 and Opus 4.8 chapters as
fallback targets for flagged requests.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015zXN1hq49epZqs3APTdCz8
…gins

Design and artifact skills now name specific styles to leave out and
extend the list when the user dislikes a choice; adhd:shape takes the
next no-input step in the same message; playwright reads the screenshot
itself for visual questions; dometrain grounding marks what no lesson
confirmed; ai-briefing's report leads with what waits on the user.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015zXN1hq49epZqs3APTdCz8
…lugins

Review findings carry file and line, why it is wrong, and how to show it
fails; PR mode leads with merge-blocking findings; fan-out audits check
each subagent's evidence before accepting it; doc-drift-detector checks
for self-contradicting numbers, dates, and names; long-run reports lead
with what waits on the user; knowledge queues the Opus 5.5 docs.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015zXN1hq49epZqs3APTdCz8
…topping point

Audit and research reports in bugs, codebase-health, mutation-testing,
discovery, architecture, and machine-health open with what waits on the
user. Dispatch briefs in codebase-health, batch-simplify, coupling,
mutation-testing, review:fanout, course-digest, and architecture:improve
say when the subagent is done and when to return early.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015zXN1hq49epZqs3APTdCz8
…tion audits

Root AGENTS.md gains a stop and keep-going rule. audit-instructions
detects think-carefully steers and reasoning-reproduction requests on
Opus 5.5 targets, vague design steers, and misplaced settled-answer
instructions; audit-prompting-postures checks finish lines, named
stops, task files, fan-out evidence checks, and report shape;
write-for-agents teaches the same when authoring. Live docs now resolve
`opus` to Opus 5.5.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015zXN1hq49epZqs3APTdCz8
…ugins

Autonomous dispatch, lane stop gates, and session-flow no longer treat a
status summary or an offer to continue as a stop; spawn specs name done
and when to stop; fan-outs check worker evidence and end in one table;
run reports lead with what waits on the user; PR prep lists only
merge-blocking findings with file, line, and repro; known-issues covers
the flagged-message model switch.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015zXN1hq49epZqs3APTdCz8
Adds the per-item provenance record, and closes the last two gaps: the
worker and merge lane launch prompts forbid ending a turn on a status
summary and lead the cycle report with escalations, and the lane-stop
alert names waiting on the operator.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015zXN1hq49epZqs3APTdCz8
PR prep leads with merge-blocking findings instead of dropping the rest;
implement-dispatch's autonomous keep-going defers to the resident-vs-clear
conditions; the AGENTS.md report headings yield to a skill's own report
shape; the loop-lane known gap now covers Opus 5.5's classifiers; the
ledger attributes G2.1 and G3.1 surfaces correctly and maps the guide's
summary items.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015zXN1hq49epZqs3APTdCz8
Its description is now 1007 codepoints, under the 1024 maximum.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015zXN1hq49epZqs3APTdCz8
@kyle-sexton
kyle-sexton force-pushed the feat/apply-opus-5-5-guide branch from 630f2b9 to f43d49e Compare September 23, 2026 12:11
@kyle-sexton
kyle-sexton marked this pull request as ready for review September 23, 2026 12:11
@chatgpt-codex-connector

chatgpt-codex-connector Bot commented Sep 23, 2026 •

Copy link
Copy Markdown

Codex Review Summary

This comment shows the latest Codex review activity on this pull request.

Review Status Commit Review trigger
📝 Code Review ✅ Completed 2026-09-23T12:21:52.488994Z f43d49e Draft marked ready
ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review" or "@codex security review".

Codex reacts with 👀 while any review is running, comments if it has suggestions, and reacts with 👍 once all reviews finish with no findings.

@claude

claude Bot commented Sep 23, 2026 •

Copy link
Copy Markdown
Contributor

Claude finished @kyle-sexton's task in 3m 12s —— View job


/review:code-review — CI code-review lane

  • Gather context and read trigger instructions
  • Invoke /review:code-review per custom instructions
  • Review changed files
  • Post summary comment

Skip-gate check: PR is open, not a draft, and the change has real review surface (168 files touched per git diff origin/main...HEAD, not the 100-file excerpt in the trigger payload) — proceeded with a full review.

Scope: this PR is almost entirely instruction/skill prose plus a handful of detection scripts. I focused the review on the logic-bearing surfaces where a prose change could silently break something machine-checked, rather than re-reading every markdown file line by line:

  • plugins/claude-config/skills/audit-instructions/scripts/{instruction-scan,emit-findings}.sh and their .test.sh files — new I8-f detector (regex, family count, decline-path counting) is internally consistent and exercised by new test cases (including the "positive think instruction" line 5 fixture, which correctly flips from not-I8-c to is-I8-f).
  • plugins/claude-config/skills/audit-instructions/reference/criteria.md, SKILL.md, context/persist-findings.md, evals/evals.json — cross-checked the "eleven scanned families / nine declined" arithmetic against the actual collect_rows call sites; it matches (I6, I8-a/b/c/f, I10, I23, I25, I27 = 9 declined; I28-a/b are the 2 emitted).
  • plugins/claude-config/skills/audit-prompting-postures/reference/postures.md + SKILL.md + evals.json — new P11 posture and the "ten"→"eleven"/generic-count rewording are consistent everywhere I checked.
  • plugins/playbooks/skills/fable-5/SKILL.md meta-rule 3 — the Opus 5.5/Fable/Opus 5 fallback-routing rewrite (biology → Opus 5, cybersecurity → Opus 4.8, Opus 5 itself refusing further biology fallback) reads as internally consistent, and matches the PR body's description.
  • plugins/autonomy/hooks/lane-stop-gate.sh — the diff is confined to the REASON=/notify strings; no control-flow change.
  • Spot-checked docs/upstream/opus-5-5-usage-guide.md's cross-references (aihero-course.md, claudedevs-cost-performance.md, docs/conventions/upstream-drift/README.md) — all exist.

No blocking or flag-worthy findings. Nothing here that a careful senior reviewer would send back — no contradicted detection logic, no broken cross-reference, no weakened gate, no unbumped plugin among the files sampled.

What I could not independently verify: this lane's allowed-tools restricts Bash to gh pr diff/view/comment/review (plus Read/Glob/Grep); running the shell test suites (instruction-scan.test.sh, emit-findings.test.sh) or scripts/check-changelog-parity.sh was refused pending approval I can't grant non-interactively. I verified the tests' logic by reading them against the script diffs (see above) rather than executing them — the PR body's own reported pass counts (97/97, 119/119, etc.) are author-claimed and unverified by me.

@claude

claude Bot commented Sep 23, 2026 •

Copy link
Copy Markdown
Contributor

Claude finished @kyle-sexton's task in 1m 17s —— View job


Security review — no issues found

  • Gather context (PR body, changed files list)
  • Understand request: custom instructions direct running /review:security-review for this PR
  • Invoke /review:security-review and follow its procedure
  • Post findings (or clean bill) to this comment

Reviewed the full diff against origin/main (git diff origin/main...HEAD, 168 files, +1283/-222) at commit f43d49e.

Skip-gate assessment: this PR has no security-relevant surface. It is almost entirely documentation and skill/agent instruction text (SKILL.md, reference/context markdown, CHANGELOG.md, plugin.json version bumps) applying an Opus 5.5 usage guide across plugins. The only non-prose files touched are:

  • plugins/autonomy/hooks/lane-stop-gate.sh — a two-line change to the human-readable REASON= string shown in the stop-gate nudge; no control-flow, input-handling, or command-execution change.
  • plugins/claude-config/skills/audit-instructions/scripts/instruction-scan.sh and emit-findings.sh — add a new I8-f regex pattern family (and widen an existing suffix class from [a-c] to [a-f]) to an offline text-scanning detector that greps local instruction files. No new external input, subprocess, or eval surface.
  • The corresponding .test.sh files — test-only additions/fixture renames for the same detector logic.

I checked the full diff for injection, secret-handling, and command-execution patterns (curl/eval/exec/subprocess/chmod/token/secret/etc.); every hit was prose documenting Claude Code configuration knobs (MAX_THINKING_TOKENS, /config, alwaysThinkingEnabled), not executable logic.

No GitHub Actions workflow files are touched by this PR, so there is nothing in zizmor's advisory lane to defer.

Conclusion: no CRITICAL, IMPORTANT, or SUGGESTION security findings. Per this lane's skip gate, no inline comments were posted.

@github-actions

Copy link
Copy Markdown
Contributor

Last security-reviewed head: f43d49efc2df3840766e729caa4cb4ca7c877494. On the next push, the relevance gate compares only the commits since this SHA; delete this comment to force a full re-review.

@github-actions

Copy link
Copy Markdown
Contributor

Claude has reviewed this PR 1 time. The lane skips further automatic reviews after 5; deleting this comment resets the count.

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: f43d49efc2

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

Comment thread plugins/docs-hygiene/skills/write-for-agents/SKILL.md Outdated
…arget model

write-for-agents removed think-carefully lines for every target, while
audit-instructions I8-f scopes that advice to Opus 5.5 because the
model-agnostic guide still recommends thinking steers. The skill now
defers to I8-f for the model the text will run on.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015zXN1hq49epZqs3APTdCz8
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants