chore(prompt-audit): close the 2026-09 prompt-audit follow-up ledger - #3939
Conversation
Codex Review SummaryThis comment shows the latest Codex review activity on this pull request.
ℹ️ About Codex in GitHubYour team has set up Codex to review pull requests in this repo. Reviews are triggered when you
Codex reacts with 👀 while any review is running, comments if it has suggestions, and reacts with 👍 once all reviews finish with no findings. |
|
Claude finished @kyle-sexton's task in 4m 11s —— View job Security review complete
This PR is the 2026-09 prompt-audit follow-up ledger closure: 346 files changed across 43 plugins, overwhelmingly prose/reference/SKILL.md edits (dating claims, removing hardcoded vendor names, rewording hook messages) plus a handful of new mechanism scripts extracted from model-executed prose per the audit's Group 1b/4 findings. I reviewed the actual diff ( No security issues found. Specifically checked:
Per this lane's scope, GitHub Actions hardening ( |
|
Last security-reviewed head: |
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: cf06fb335c
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
0bc444b to
9b9c45d
Compare
|
Claude has reviewed this PR 2 times. The lane skips further automatic reviews after 5; deleting this comment resets the count. |
9b9c45d to
f492fe0
Compare
Pull request was converted to draft
|
Claude finished @kyle-sexton's task in 3m 41s —— View job CI code review complete
No high-signal findings. Skip-gate checks passed (PR is open, not a draft, and has real review surface), so I did a full pass per the Given the size (347 files vs. Verified, not just read:
Nothing here rises to a blockable finding — the two real defects a prior automated pass (Codex) found were both already fixed and independently re-verified by the author with reproduction steps, and I re-confirmed both fixes are present and correct at head Per this lane's scope, security-relevant surface (injection, trust boundaries) was already covered by the security-review lane's clean pass on this branch and is not re-litigated here. Branch: |
|
Claude finished @kyle-sexton's task in 1m 20s —— View job Security review complete
Last security-reviewed head was No security issues found. This delta is exactly the fix-up for the two
No injection, auth, secret-handling, or trust-boundary surface in this delta. Per this lane's scope, GitHub Actions hardening is zizmor's advisory lane and this delta touches no workflow file. |
68c7ffc to
f83d708
Compare
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: 68c7ffcee8
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
aaffcd3 to
086665f
Compare
Same tree as the reviewed follow-up head plus the calendar and keyword idempotency pins. Collapsed to one cursoragent-signed commit so GitHub's required-signatures rule can evaluate the squash: unsigned and unknown_key commits on the previous head blocked merge even though GitHub would sign the squash result. Co-authored-by: Kyle Sexton <153232337+kyle-sexton@users.noreply.github.com> Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
086665f to
2f5d827
Compare
…opped The description was still the fork's text, so the Windows kernel-object census, the per-session hook event log and its report, the setup action's skill-usage and hook-log roots, and the sink's per-session routing were all absent from the generated catalog. Takes main's wording verbatim and keeps the branch version. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
exit 0 records 0ms on CI, and ratio.py fail-closes when every comparison-arm sample is zero milliseconds. sleep 0.01 still does not drain stdin, so the 141-fabrication assertion still holds. Co-authored-by: Kyle Sexton <153232337+kyle-sexton@users.noreply.github.com>

No related issue: the follow-up ledger of the prompt-audit sweep is tracked by the record
docs/specs/prompt-audit-skills-2026-09.md, not by an issue.Summary
Closes the follow-up ledger of the 2026-09 prompt-audit sweep (#3770): every item F2 to F24 in
docs/specs/prompt-audit-skills-2026-09.md## Follow-upsnow names its closing commit (20 lines prefixed "Closed"), F3 and F14 are recorded as not PR deliverables with owners, and F25 is inventoried. 43 plugins change across 121 commits plus 4 merge commits fromorigin/main. The small defects are fixed directly, the mechanism changes ship as scripts with tests, the verification debt is stamped with dates and recheck triggers, and the surfaces the sweep did not audit (plugin-levelreference/trees, hooks prompt text, the output style,.claude/rules,CLAUDE.md,AGENTS.md) were audited with the same method and applied the same way. Reports and decisions per follow-up are under the author's.work/prompt-audit-follow-ups/and are summarized in the record's## Follow-up PRsection.Fix
(prompt-audit follow-up F<n>).work-items/scripts/lane-telemetry-upsert.sh(F11), the syncedcompose-statusline-wiring.shin context-guard and rate-limit-guard (F15), context-guard setup's pre-computed probes (F17), the syncedplugin-quality/scripts/context-zone.shwithsync-context-zone.sh --checkand--check-bumpwired into CI's plugin-gate job (F19),mutation-testing/scripts/suppression-lint.sh(F21), pull-request Gate 5 parameterized on the discovered reviewer (F7), raw-URL cross-plugin citations and a link-free convention name in autonomy (F8).list-corpus, knowledge's fence gate console encoding, education's teach workspace across worktrees) and several Windows-only test vacuums (a CRLF from jq that made checks pass without inspecting anything). One residual is verified on CI's Linux lane: repo-fleet-hygieneaudit-fleet's 33 load-sensitive GitHub-evidence cases..claude/rules, the renderedAGENTS.mdrules table, the one output style, and the user-facing strings of 70 hook scripts audited and applied; nohooks.jsondeclares a prompt-type hook.## Follow-up PRanddocs/topics/prompt-audit-follow-ups/was pruned.origin/main's current version for the plugins main overtook; merge commits compose both sides of every content conflict (named in each merge commit message).Verification
scripts/check-changelog-parity.shmodes (--check,--check-bump origin/main,--check-preserved origin/main,--check-order): pass at the branch tip.check-purged-em-dashes.sh,check-skill-precompute-compose.sh --all,render-index.sh check --file AGENTS.md,validate-plugin-contracts.mjs,sync-context-zone.sh --checkand--check-bump origin/main,check-shell-portability.sh origin/main,check-docs-only-gate.sh --check,check-skill-count-claims.sh --check,check-contract-slice-prune.sh --check-diff origin/main,typos, repo-wide markdownlint: pass.check-cross-plugin-source-drift.sh --checkreports the same 20 unregistered clusters on a cleanorigin/maincheckout on the author's Windows host, so that gate's verdict is CI's Linux lane.scripts/affected-tests.sh --runcould not finish on the author's host under load; every suite each follow-up touched was run directly and is green there, and each dispatched change was checked by a fresh-context verifier against the diff and the run before commit (per-cluster reports name the counts). CI runs the whole corpus.check-skill.shon every touched skill,check-skill-precompute-compose.sh --paths,check-evals-quality.shwhere evals changed, shellcheck and the portability check on every script, markdownlint, typos, and theai-slopdetector on every touched file.Related
docs/specs/prompt-audit-skills-2026-09.md(## Follow-ups,## Follow-up PR).fe3ada293f429da5416d89aa11555f4082222416(gh api "repos/{owner}/{repo}/contents/docs/topics/prompt-audit-follow-ups/PLAN.md?ref=fe3ada293f429da5416d89aa11555f4082222416" --jq .size)..work/prompt-audit-follow-ups/reports/F2-playbooks-reference-applied.mdfor the playbooks maintainers to file through/playbooks:update; filing public issues was outside this PR's ask.babysit-readiness-gate.shbadge grammar) is inventoried in the record, not delivered here..claude/rules/pr-body-contract.mdis the one rule with no YAML frontmatter. That file is sync-manifest-managed bymelodic-software/standards(CI's managed-files-guard refuses direct edits), so the frontmatter belongs in that repository and lands through standards-sync.Follow-ups (verbatim from the record)
Inventoried here as they arise and shipped in the PR body verbatim.
.claude/rulesand theAGENTS.mdrow they render), b618ed3 (the one output style), c93b18c, 2ba9047, and fc0d3b7 (hooks prompt text in guardrails, disk-hygiene, and autonomy); nohooks.jsondeclares a prompt-type hook, so the hooks surface was the user-facing strings of 70 hook scripts, 65 of them clean. Reports and decisions for both surfaces sit under.work/prompt-audit-follow-ups/beside the six reference-tree ones. In progress: the six plugin-level reference trees are audited and applied (performance 61d3eb9 and d0ce988; autonomy 67e8244; architecture 1ca9d3c; coupling cdcee0e and review 075fc0a for the same topic-docs pointer regression; playbooks 4e2a0f4; context-guard 4906eed and 3452edf; rate-limit-guard b4264ea), with reports and decisions under.work/prompt-audit-follow-ups/; hooks prompt text, output styles,.claude/rules,CLAUDE.md, andAGENTS.mdare still to audit. Audit the out-of-scope prompt surfaces the same way: hooks prompt text, output styles,.claude/rules,CLAUDE.md,AGENTS.md, and the plugin-levelreference/trees that skills load on invocation (autonomy,architecture,performance,playbooks,rate-limit-guard,context-guard); the performance auditor notes thatsnapshotandverifyboth mandate readingplugins/performance/reference/harness-integrity.md, which likely mirrors the archaeology the skill bodies shed.claude-config:unhobbleon demand, and the operator owns when to run it. Behavior measurement beyond the wave-1 spot-check: route toclaude-config:unhobble.docs/topics/prompt-audit-skills/PLAN.mdgraduated into Brief and the slice was pruned (contract-slice prune gate); the PR body names the pre-prune commit.plugins/skill-quality/scripts/check-skill.shcheck 3 hard-fails any trigger phrase dropped versus the base ref. That blocks prompt-audit's documented fix for trigger-case enumeration (near-synonym lists become intent categories). Change check 3 to a warning, update its tests, and record the deliberately dropped phrases per skill in this record. Must land before the PR so the skill-quality CI gate passes. Landed in a694011; two out-of-scope surfaces still describe check 3 as a hard-FAIL gate and should follow: the comment atplugins/docs-hygiene/skills/audit-noise/scripts/emit-findings.sh:220and the eval fixtureplugins/docs-hygiene/skills/audit-noise/evals/fixtures/negation-trigger-fence.md:9(the docs-hygiene commit corrected the same script's header comment and its eval case 12).skills/worktree/reference/gather-block.md, "The pre-compute block runs as one shell invocation"); G2 e564b9b (disk-hygiene), bf70d79 (performance), 4e0e4d8 (rate-limit-guard), 3ee8271 (work-items); G3 203fb36 (claude-memory), c0275b4 (claude-config, which owns thepre-v2.1.211record), 7659e32 (claude-ops); G4 5404ba7 (context-budget), 4c792b9 (context-guard), 4532e37 (plugin-quality), 177b179 (skill-quality), 993ce43 (instruction-placement); G5 0d49296 (computer-use), f344689 (overengineering), 875b8da (improvement), 70dfe3e (code-tidying), 69a8819 (codebase-health), 1aa1dcf (architecture), caaa69c (mcp-tools), 22d20a4 (docs-hygiene), 1fd6561 (repo-hygiene); G6 da9b14a (autonomy, basis by class), 809617c and 32ebdee (discovery, six records inreference/parent-contract.md), 0bb93b9 (discipline), 577c359 (testing), 16a2604 (playbooks), 634ef1b (review), 69913c1 (planning). Across 122 claim rows: 80 stamped, 32 corrected and stamped, 9 removed as unverifiable from any reachable source, 1 needing no edit; 36 skill sites outside source-control now point at the gather-block record. ReportsF6-G1.mdtoF6-G6.mdunder.work/prompt-audit-follow-ups/reports/carry the per-claim basis and recheck trigger. Verify and stamp the undated harness-behavior claims the audit flagged asI12items (session-flow:/recaptrigger and the Skill-invocable allowlist, the usage-limit reset surface,cleanupPeriodDaysdefault,/cleartranscript and scheduled-task behavior). Each becomes a four-part upstream-drift record or a doc pointer. Collected per plugin as the waves run. source-control adds: GitHubmergeStateStatusprecedence andbaseRefOidstaleness (freshness.md), the permission-mode and wrapper-strip claims in safety.md, and the ScheduleWakeup clamp,/loopexpiry, Monitor-on-resume, and sandboxed-GraphQL claims across babysit-prs, babysit-loop, pull-request, and worktree. disk-hygiene adds four undated harness-version claims (safety-model.md:262,:313,:315on 2.1.207 and 2.1.218pluginConfigsscope and PowerShell hook firing;clean/SKILL.md:382on the v2.1.211 auto-mode prompt). performance adds the undated benchstat-delta-testclaim atsnapshot/SKILL.md:94-96, andplugins/performance/reference/harness-integrity.md(out of audit scope, mandated reading forsnapshotandverify) likely carries the run archaeology the four skill bodies shed (see F2).reference/reviewer-shapes.mdas dated records;scripts/babysit-readiness-gate.shstill couples to one vendor's badge grammar and is its own follow-up).pull-requesthardcodes one vendor's review bot (login, emoji signalling, timing) inreference/monitor.mdgotchas andreference/readiness.mdGate 5, against the file's own "discover actors, don't hardcode them" rule. Parameterize Gate 5 on the discovered reviewer login and move the vendor shapes into a dated reference-shapes note with a recheck trigger.source-controlcites sibling-plugin files by relative path (../../../../autonomy/...,../../../../../prompts/...) inbabysit-loop/reference/promotion-evidence-resolution.mdandbabysit-prs/reference/safety.md. Those resolve only in the marketplace checkout, never in an installed plugin. Convert to the raw-URL form the plugin already uses atbabysit-loop/SKILL.md:40. review adds:agents/ecosystem-specialist.md:22,fanout/context/fix-pass-mode.md:7,quality-gate/context/close-out.md(four sites), andquality-gate/context/spec.md(three sites) cite marketplacedocs/paths or sibling-plugin files by relative path. review also adds two undated claims to F6: the bundled/code-reviewand managed Code Review service tiers inquality-gate/context/pr.mdandcode.md, and the "built-in/security-reviewis unusable in CI" claim insecurity-review/SKILL.md. The autonomy applier converted six of that plugin's own convention and sibling citations to marketplace URLs; those were reverted to the relative form before the PR becausescripts/validate-plugin-contracts.mjsforbids the autonomy plugin naming the org or a vendor, so the autonomy sites stay under this follow-up.planning/skills/wayfind/SKILL.mdpre-compute silently coerces a non-stringcontainer_labeltowork-map, whichcontext/tracker-mechanics.mdsays is a configuration error that must never proceed. Make the pre-compute fail loud or surface the raw value, matching the doc.work-itemscarries the same 60-line lane-telemetry upsert as a fenced shell block inwork-loop/reference/telemetry-upsert.mdandattend-queue/reference/telemetry-upsert.md, transcribed by the model on every cycle, with a classifier-fallback section asking it to re-derive gate order by hand (prompt-audit Group 4, an LLM executor for a deterministic plan). Extract it intoplugins/work-items/scripts/lane-telemetry-upsert.shwith a co-located test, taking lane, instance, repo, issue, and body-file arguments and exiting non-zero on each refusal branch; both references then invoke it. Deferred from the audit because it is a mechanism change, not a prose hunk. work-items also adds six undated harness andghclaims to F6 (classifier refusals ofpermissions.allowwidening and of thereclaimcall, the compound-shell block, sandboxed GraphQL 403).shell: bashfrontmatter selects the shell for!...`` injections. Skills whose pre-compute block became empty when the git lines moved into body calls still carry the key inert (debugging F5 found one). Sweep every SKILL.md: where no injection remains, drop the key;check-skill.shcheck 19 stays green either way.playbooks:borispresents Fable 5 as the current top model and its launch-era classifier behavior as current (skills/boris/SKILL.md:58,133); upstream has not published Fable 5.1 tips. Re-sync through/playbooks:updatewhen it does, and until then qualify the Model row "as of the 2026-07-24 sync"./playbooks:update; thefable-5-1.mdchapter exists, the F8 tail closed under F8, and the F6 tail closes under F6. Thefable-5playbook's own regeneration trigger ("a model-version change",skills/update/SKILL.md:39) has fired with Fable 5.1. The audit adds the guide-backed minimum, afable-5-1.mdadaptation chapter; regenerating the whole pack from Fable 5.1 is the maintainers' larger call. playbooks also adds to F8:reference/model-adaptation/opus-5.md:207-208cites a probe record (thinking-off-probe-2026-07-26.md) that exists nowhere in the repository. And to F6: the cache-pricing stamp atskills/fable-5/context/orchestration.md:97carries a date but no recheck trigger.unwrap-before-compose.md(synced betweencontext-guardandrate-limit-guard) is a pure function of the effectivestatusLinestring that the model hand-executes over roughly a hundred lines of prose, with eight eval cases checking the arithmetic (prompt-audit Group 4). Extract it into a syncedscripts/compose-statusline-wiring.shwith the round-trip check inside, shrink the reference to the contract, and turn those eval cases into script tests. Deferred from the audit as a mechanism change. rate-limit-guard also adds to F6: the undated "Monitors is an experimental Claude Code component" claim inreference/reader-contract.md:206-209.claude-ops/skills/plugins/SKILL.md:268-276records that its own probe's recheck trigger has fired (the CLI moved from 2.1.218 to 2.1.240 with the claim un-retested). Re-run the probe and refresh the stamp. claude-ops also adds nine undated harness and upstream-issue claims to F6 (bundleddoctorgating,audit-native-overlapalias examples,inventorycommand aliases, the WebFetch truncation window, theCLAUDE_PLUGIN_DATAexport claim, thelanes"verified on this machine" lines, theobservabilitysession_idand Stop-hook gotchas, upstream issue states inread-routing.mdandsync.md, and the triggerlesssurfaces.mdstamp) and two measured figures (backups/retention, the 97 percent and 50 MB figures inobservability).plugins/repo-fleet-hygiene/skills/audit/scripts/audit-fleet.test.shfails 5 of 180 cases on this host (the worktree-root-unconfigured placement and header cases, the symlink discovery-root case, the intermediate-symlink case, and the unreadable discovery-root case); the scripts are untouched by the repo-fleet-hygiene commit and the finding-kind table assertion passes.${CLAUDE_SESSION_ID}must be visible to the consumer, per the reader contract).context-guard/skills/setup/SKILL.mdruns four fixed read-only probes (jq presence, installed shim versus shipped source, session snapshot,zones.json) as model-issued Bash calls where a## Pre-computed contextblock would run them before the body loads (prompt-audit Group 4). Adding one is a mechanism change: the block must passscripts/check-skill-precompute-compose.shand stay inside the worktree guard's rule that a composed block expands nothing but bare$HOME, so it is deferred from the audit. context-guard also adds to F6: the undateddisableAllHooks/allowManagedHooksOnlyclaims inskills/setup/SKILL.md:93-96andreference/reader-contract.md:503-507, the undated PowerShell routing note instatusline-edit.md:106-109, and the folklore-number paragraph atreader-contract.md:383-391, which is dated but has no recheck trigger.autonomy/reference/autonomous-pipeline-reminder.md(out of audit scope; cited only by the README and a hook) rewords the vendor's autonomy block under the repo's no-copy rule and omits the Fable 5.1 clause "Do not stop because the context or session is long"; the guide calls the opening sentence load-bearing as written. Weigh the no-copy rule against that claim and add the missing clause in the plugin's own words. autonomy also adds to F6: the undatedAGENTS.md-reachability claim stated three times (skills/setup/SKILL.md:267,context/prerequisite-resolution-slice.md:38-39,reference/prerequisite-resolution.md:86-88), the undated empirical telemetry claims inreference/telemetry.md, and the "shipped first-party mechanisms today" claims inreference/runner/escalation.md:140-152.sync-context-zone.sh --checklands in this PR's closing commits).plugin-quality/skills/audit/SKILL.md:58-92has the model resolve the context zone by hand from inlined band tables, a staleness window, a version floor, and a combination rule thatplugins/context-guard/scripts/context-zone.shalready implements (prompt-audit Group 1b and Group 4). Ship a byte-identical synced copy atplugins/plugin-quality/scripts/context-zone.shwith its test, register it inscripts/cross-plugin-source-registry.txtwith async-context-zone.sh --checkentry, and have the gate andsetup/SKILL.md:28-30call it. Deferred from the audit as a mechanism change. plugin-quality also adds to F6: two live doc-page titles quoted undated inagents/auditor.md:117-119, thecontext: forkand cloud-scoping claims inreferences/component-types/skill.md:18-24, and six dated stamps with no recheck trigger. skill-quality adds to F6: three undated harness claims outside the dated stamp incheck/SKILL.md:160-172, and thesetup/SKILL.md:16-20stamp that has no recheck trigger. instruction-placement adds to F6: the undated "other agents resolve nearest-wins" claim inrealign/context/apply-recipes.md:95-97. context-budget adds to F6: thev2.1.232measurement ataudit/SKILL.md:226-228, the/doctoravailability anddisableModelInvocationclaim ataudit/SKILL.md:34-36, the cited-but-undated mechanism claims inaudit/reference/engine.md:25-30with the dangling "verified version" referent at:52-53, and the wall-clock range ataudit/SKILL.md:93. computer-use adds to F6: the dated surface table indiagnose/SKILL.md:62-63and the dated basis indiagnose/reference/windows-quirks.md:5-6, both without a recheck trigger. overengineering adds to F6: the undated harness-behavior claim in the gather blocks of all three skills (audit/SKILL.md:20-23,delta/SKILL.md:19-23,realign/SKILL.md:19-22, covered by the one dated record the worktree skill will own) and the undated/loopcapability claims indelta/context/recurring-wiring.md:37-38,51-53. improvement adds to F6: four undated GitHub REST and Claude Code CLI claims infind/context/ci-health.md:32-41,find/SKILL.md:235-237, andfind/context/unattended.md:74-75. docs-hygiene adds to F6: the bundled/batchskill claim inextract-ssot/actions/batch.md:35,281, four undated external benchmark figures acrossextract-ssot/SKILL.md:27,context/anti-patterns.md:129, andcontext/decision-framework.md:27-59, and the undated upstream-publishing claim inaudit-encapsulation/context/public-surface-contract.md:5. code-tidying adds to F6: the CodeScene agentic-refactoring figure intidy/reference/scope-budget.md"Research lineage" has no resolvable source; the audit dropped the number and kept the qualitative claim until a publication URL and read date are recorded. repo-hygiene adds to F6: the sourced-but-undated${CLAUDE_SKILL_DIR}substitution-scope claim inclean/reference/invocation-forms.md. disk-hygiene adds to F6: four undated harness-version claims acrossclean/SKILL.mdandclean/reference/safety-model.md(report F15). codebase-health adds to F6: the undated harness-capability claim ataudit/SKILL.md:25-28, verified true by the auditor on 2026-09-04 and needing only its dated record. architecture adds to F6: the undated pre-compute execution claim atimprove/SKILL.md:25-28. mcp-tools adds to F6: three cited-but-undated Claude Code client-behavior values inaudit/reference/checklist.md:38,105,106. performance adds to F6: the undated benchstat flag-set claim insnapshot/SKILL.md:94-96.provenance/skills/auditspells one tier two ways:not-foundinSKILL.md:2,82,227andsource-not-identifiedinreference/rubric.md:297, andscripts/emit-findings.shwith its test asserts both. Pick one spelling, change the script andemit-findings.test.shwith it, and align the markdown in the same commit. Deferred from the audit because the fix crosses into a script and its suite.mutation-testing/skills/setup/SKILL.md:80-89has the model re-derive a suppression entry'sfinding_idhash from its constituents and check node-kind membership by hand (prompt-audit Group 1b and Group 4, the same shape as F19). Shipplugins/mutation-testing/scripts/suppression-lint.shwith a test implementing the two published derivations and the membership check, have setup call it, and retarget setup eval 5 and audit eval 3 from "the model re-derives" to the script. Deferred from the audit as a mechanism change.plugins/ai-briefing/skills/setup/evals/evals.json:33prompts/ai-briefing:setup --with-build-deps, but the skill's contract isapply install-build-deps; the case exercises a flag the skill does not accept. Retarget the prompt to the contract form. Observed by the ai-briefing auditor outside the audit's markdown scope.plugins/dometrain/skills/sync/context/update.mddocuments the maintainer-only--refresh-baselinecommand through${CLAUDE_PLUGIN_ROOT}, which resolves to the installed plugin cache in a normal session, while the next paragraph forbids running it anywhere but a working clone; the script writes next to itself either way. Give the command as a clone-relative path, or document that the flag is only safe under--plugin-dir. A script-safety contradiction, not a prose hunk; observed by the dometrain auditor.plugins/claude-config/skills/audit-instructions/reference/criteria.mdfor each recurring shape in Catalog gaps: dated stamps with no recheck trigger, migration-relative phrasing inside reference and context files, routing text that names a skill absent fromplugins/, sibling-file meta-commentary, and maintainer rationale inside model-facing YAML comments; the rest are one-offs and stay listed.cleansuites) and cleared it; the autonomylane-stop-gateFIFO case passes on this host unchanged. One residual stays, not fixable on this host: repo-fleet-hygieneaudit-fleet's 33 load-sensitive GitHub-evidence cases, verified on CI's Linux lane. ReportsF10-C2.mdandF10-C5.mdunder.work/prompt-audit-follow-ups/reports/. The status below is as recorded on 2026-09-07 before C2 landed. In progress: clusters C1, C3, and C4 landed (C1 1c75e31 and f385946 provenance, 037552c cloud-bootstrap; C3 eab3540 claude-ops, eec9018 claude-config; C4 6d34a94 knowledge, 09c8d00 education, 944e0cb disk-hygiene, e07b89d repo-fleet-hygiene). Three of the recorded symptoms were real cross-platform defects, not host quirks: provenancelist-corpussubtracted the corpus root as a string, knowledge's fence gate wrote to a cp1252 console, and education's teach workspace split across a repo's worktrees. Four recorded symptoms did not reproduce (docs-hygieneemit-findings, claude-configemit-findingsandpermission-rule-check6b, claude-opsfleet-state, the last resolved upstream by main's rewrite). Still open: cluster C2 (work-itemsgenerate-adaptercase 116, planninginterview-defensesdigests, instruction-placementverify-load), the autonomylane-stop-gateFIFO case held back from C4 while a verifier held that plugin, theinterview-defensesfrontmatter digest that F12'sshell: bashremoval moved and that needs re-pinning, repo-fleet-hygieneaudit-fleetat 33 load-sensitive GitHub-evidence failures the C4 report characterizes, and one uncaptured error in repo-hygienecleancheck 7 with main ahead of the branch in those scripts. Not an audit finding, recorded so it is not mistaken for one:.claude/hooks/cloud-bootstrap-plugins.test.shfails 15 of 32 assertions on this Windows host ("not installed at user scope") with.claude/cloud-bootstrap.shand the suite byte-identical toorigin/main. The failure is environmental or pre-existing; confirm on CI and file separately if it reproduces there. Same status forplugins/docs-hygiene/skills/audit-noise/scripts/emit-findings.test.sh("tier is looked up as IMPORTANT", "Location is repo-relative") andplugins/provenance/skills/audit/scripts/list-corpus.test.shandemit-findings.test.sh("a directory target lists its markdown"), which fail on this host with their scripts and suites byte-identical toorigin/main. Same again forplugins/work-items/skills/onboard-adapter/scripts/generate-adapter.test.shcase 116, and for the nine eval-case digest assertions inplugins/planning/tests/interview-defenses.test.sh(interview/evals/evals.jsonunchanged since the digests were pinned; local jq 1.8.2), and for four Windows temp-path cases inplugins/instruction-placement/scripts/verify-load.test.sh(selected by a basename collision ontypescript.md; the probe and suite are unchanged on this branch), and forplugins/claude-ops/skills/audit-install-state/scripts/install_state.test.sh(a Windows filename-syntax error on a fixture path) andplugins/claude-ops/skills/audit-skill-visibility/scripts/audit_skill_visibility.test.sh(noinstalled_plugins.jsonin the temp config), both with scripts and suites byte-identical to HEAD, and forplugins/claude-ops/skills/plugins/scripts/fleet-state.test.sh, which fails a varying subset of its 74 cases on this host (six inside a check-skill run, two when run alone) with the scripts byte-identical toorigin/main. Same again forplugins/claude-config/skills/audit-instructions/scripts/restatement-scan.test.sh(two I29 fixture cases, script and fixtures byte-identical toorigin/main) and the oneemit-findings.test.shcase downstream of it ("Action names a body cut"), which reads the same scanner's output. Same again forplugins/claude-config/skills/audit-permission-grants/scripts/permission-rule-check.test.shcase 6b ("vendored copies excluded, exactly one finding"), whose script and suite no branch commit touched (main has since tidied the suite in ac7eeea). The fleet gather block itself ("the harness runs a skill's whole pre-compute block as one shell invocation") is an undated harness claim in about 55 skills; one dated four-part record on the worktree skill, which owns the mechanism, with the copies pointing at it, clears every site at once. discovery adds six undated claim families across thirteen files (silent preload failure,AskUserQuestionand plan-mode tools filtered from non-fork subagents, the Workflow tool absent from subagents, background as the default execution mode, spawns permission-classified before launch); the fix is one dated record per claim in the plugin'sreference/parent-contract.mdwith the skills pointing at it. claude-config adds the undatedpre-v2.1.211boundary at six body sites (the dated owner isaudit-permission-state/reference/criteria.md), dated-but-triggerless stamps across eight files, theconflict-scan.shprecision figures inconflict-criteria.md, and the "Fable 5 subpage" pointers inaudit-prompting-postures/reference/postures.mdthat need a Fable 5.1 sibling once it exists. discipline adds five files of undated fork-mode harness claims (sweep-all/SKILL.md, its two references,scrutinize-dont-coast/SKILL.md,use-your-skills/SKILL.md). claude-memory adds the undated upstream-issue state ataudit/reference/official-guidance.md:168. testing adds the xUnit v3 and .NET 10 framework-trap claims (diagnose/SKILL.md:68,diagnose/context/investigate.md:16,write/SKILL.md:74) and theplaywright-cliversion floor inrun-e2e/context/e2e.md:12. planning also adds two undated harness claims to F6: the agent-teams "experimental, default-off" status inplan/SKILL.mdand the "cannot read effort or advisor state" claim ininterview/context/session-config.md.plugins/ai-slop/skills/audit/scripts/detect.test.shfails its four "git absent" cases (4 of 202) on this Windows host because the test symlinks the shell builtinprintfinto a fake PATH directory (ln: failed to create symbolic link); the scripts are unchanged by the ai-slop commit.plugins/disk-hygiene/skills/clean/scripts/guard_launch_monitor.test.shfails its two telemetry-sink cases andhygiene.test.shfailstest_stash_must_exist_in_an_independent_checkoutandtest_preview_allows_root_children_os_managed_snapshoton this host with the scripts byte-identical to HEAD; neither case reads markdown, and the frontmatter-belt assertions intest_hygiene.pythat do readclean/SKILL.mdpass after the disk-hygiene commit.plugins/code-tidying/skills/audit-comment-residue/scripts/detect.test.shfails 4 of 53 cases on this host (embedded quote and backslash unescaping in the preview script, arrow-in-filename, tab-bearing path) withdetect.shand the suite byte-identical to HEAD; the code-tidying commit touches only prose the script does not read.plugins/knowledge/skills/docpage-digest/scripts/check-fences-exact.test.shfails six cases on this host with aUnicodeEncodeErrorwriting U+2265 to the cp1252 console, script and suite byte-identical to HEAD; the knowledge commit touches no fence, quote payload, or script.plugins/education/skills/teach/scripts/list-workspaces.test.shfails "worktree lists the MAIN repo's workspace" on this worktree checkout with the script unchanged.plugins/source-control/scripts/babysit-readiness-gate.shmatchesP[0-3]severity badges, one vendor's finding-format grammar, the same coupling F7 removed frompull-requestone skill over (it servesbabysit-prs). Parameterize the badge grammar on the discovered reviewer shape thatpull-request/reference/reviewer-shapes.mdrecords, with the script's test updated in the same commit. A script change outside the F7 commit's blast radius (from the F7 applied report).PLAN.md (graduated into the record and pruned; last held at fe3ada2)
prompt-audit-follow-ups
Brief
TLDR
Close every open follow-up the 2026-09 prompt-audit record inventories (F2 to F24, less the two already done) in one follow-up PR on top of the merged sweep: the small defects directly, the mechanism changes as scripts with tests, the verification debt as dated records, and the unaudited surfaces with the same audit method.
Goal
docs/specs/prompt-audit-skills-2026-09.md## Follow-upsreads as a closed ledger: each of F2 to F24 either names the commit that closed it or states, in one sentence, why it stays open and who owns it. Every skill that loads areference/tree, hook prompt, output style, or rule has had those surfaces audited against Claude Fable 5.1. Every deterministic procedure the audit found the model transcribing by hand runs as a script with a test. Every undated harness or upstream claim the audit withheld is stamped with a date and a recheck trigger, pointed at a dated owner, or removed.Constraints
chore/prompt-audit-follow-upsfrommainat f62c1b3, one worktree, one PR. One commit per follow-up (or per plugin inside a follow-up that spans plugins), each naming the follow-up id in its subject.origin/main's current version at commit time and are renumbered once before the PR if main overtakes them.check-skill.shon every touched skill,check-skill-precompute-compose.sh --paths,check-evals-quality.shwhere evals change, markdownlint, typos, theai-slopdetector; scripts get shellcheck and a co-located test that runs green on CI's Linux lane.scripts/validate-plugin-contracts.mjs); its citations get a link-free convention name, never a marketplace URL.general-purpose; at most three concurrent; every dispatch held while~/.claude/rate-limit-guard/rate-limits.jsonshowsfive_hour.used_percentageabove 85.worktree.baseRefin any settings file (.claude/rules/worktree-base-ref.md).Acceptance criteria
Direct fixes, lead:
emit-findings.shcomment and thenegation-trigger-fence.mdfixture no longer describe check 3 as a hard-FAIL gate.container_label, matchingcontext/tracker-mechanics.md.SKILL.mdcarries ashell:frontmatter key with no!`injection left in the file;check-skill.shcheck 19 stays green.autonomous-pipeline-reminder.mdcarries the "do not stop because the context or session is long" clause in the plugin's own words.SKILL.md,reference/rubric.md,emit-findings.sh, and its test.apply install-build-deps.--refresh-baselineis documented as a clone-relative command or as safe only under--plugin-dir.criteria.mdgains one row per recurring catalog-gap shape (dated stamp without recheck trigger; migration-relative phrasing in reference and context files; routing text naming a skill absent fromplugins/; sibling-file meta-commentary; maintainer rationale in model-facing YAML comments).Mechanism changes, one dispatched implementer each, script plus test plus the doc wiring:
plugins/work-items/scripts/lane-telemetry-upsert.shwith a test; both telemetry-upsert references invoke it.scripts/compose-statusline-wiring.shin context-guard and rate-limit-guard with the round-trip check inside;unwrap-before-compose.mdshrinks to the contract; the arithmetic eval cases become script tests.## Pre-computed contextblock that passescheck-skill-precompute-compose.shand the worktree guard's$HOME-only expansion rule.plugins/plugin-quality/scripts/context-zone.shwith test, registry entry, andsync-context-zone.sh --check; the audit gate andsetup/SKILL.mdcall it.plugins/mutation-testing/scripts/suppression-lint.shwith a test; setup calls it; setup eval 5 and audit eval 3 retarget to the script.docs/relative citation in source-control, review, and playbooks converted to the raw-URL form; the nonexistent probe record citation inopus-5.mdremoved or pointed at a real file; autonomy's six sites given a link-free convention name.Audit and verification, dispatched per surface or per plugin:
reference/trees of autonomy, architecture, performance, playbooks, rate-limit-guard, and context-guard, plus hooks prompt text, output styles,.claude/rules,CLAUDE.md, andAGENTS.md, audited with the same brief and applied the same way; reports under.work/prompt-audit-follow-ups/reports/.scripts/affected-tests.sh --runover this branch's diff reports no unexplained failure.Closing:
/claude-config:unhobbleon demand.## Follow-upsnames the closing commit per item; this Brief graduates into the record as a## Follow-up PRsection and the slice is pruned; the PR body carries the closed ledger and the pre-prune commit.check-changelog-parity.shmodes,check-purged-em-dashes.sh, repo-wide markdownlint and typos,check-skill-precompute-compose.sh --all,render-index.sh check --file AGENTS.md,validate-plugin-contracts.mjs,check-contract-slice-prune.sh --check-diff origin/main, and CI on the PR.Captured assumptions
claude-apiprompt-audit guide is unchanged since the sweep.Out-of-scope
Deferred questions
Plan
Phase A (lead, direct): F22, F23, F9, F20, F13, F18, F5 residual, F24, F12, F16 bookkeeping.
Phase B (dispatched implementers, three at a time): F11, F15, F21, then F19, F17, F7, F8.
Phase C (dispatched auditors and verifiers): F2 surfaces, F6 per plugin group, F10 per suite cluster.
Phase D (lead): record ledger, version renumber against current main, Brief graduation and prune, gates, PR.
🤖 Generated with Claude Code
https://claude.ai/code/session_01S1Vu8qA5MKiSnXrWMEdbFW