docs: replace in-place skill-frontmatter restatements with upstream references - #3524
Conversation
…eferences Frontmatter-alignment sweep against a live fetch of the official skills frontmatter reference (code.claude.com/docs/en/skills, 2026-08-31): - Convert the 11 bare restatements found by the verified census to pointers at the official reference, or to four-part records (claim, basis, as-of date, recheck trigger) where the surface needs the fact inline to function. - Upgrade date-only stamps (invocation-mode, permission-rule-hygiene, MIGRATION-PLAYBOOK, naming sources, claude-config audit references) to the full four-part shape, and refresh the stale aggregate listing reading (117,695 -> 130,470 chars, 2026-08-31 instrument run). - Add four-part records to the three constant-encoding scripts (check-skill.sh, check-listing-budget.sh, audit_skill_visibility.py). - Re-anchor the pre-v2.1.216 name-history claim to the changelog basis in PLUGIN-PHILOSOPHY and the skill-authoring playbook; the skills page no longer pins that history and now records the v2.1.216-245 prefix-doubling quirk fixed in v2.1.246. - Declare user-invocable: true on the 5 skills missing it (explicit-key posture; documented default, no behavior change). - Record the fired no-gate recheck trigger in the upstream-drift CHANGELOG (1.6.2): decision re-derived and upheld. - Commit the task contract (docs/topics/skill-frontmatter-alignment/PLAN.md). Validation: affected-tests 140 shell suites pass; the 7 selected Python suites pass in their own lane (562 tests); check-skill.sh passes on all 8 touched skills; markdownlint, typos, editorconfig-checker, shellcheck, shfmt, and the pinned ruff wrapper are clean. The machine-health PowerShell suite was selected but cannot run on this Linux runner; that plugin's change is one frontmatter line. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_015DrfSCgSk56RoVv2gKWmQm
Codex Review SummaryThis comment shows the latest Codex review activity on this pull request.
ℹ️ About Codex in GitHubYour team has set up Codex to review pull requests in this repo. Reviews are triggered when you
Codex reacts with 👀 while any review is running, comments if it has suggestions, and reacts with 👍 once all reviews finish with no findings. |
|
Claude finished @kyle-sexton's task in 3m 10s —— View job Code Review (
|
|
Claude finished @kyle-sexton's task in 55s —— View job Security review —
|
|
Last security-reviewed head: |
|
Claude has reviewed this PR 2 times. The lane skips further automatic reviews after 5; deleting this comment resets the count. |
…ntion The docs/topics/skill-frontmatter-alignment/ slice is Contract tier: committed on the task branch while the work ran, pruned before merge (contract-slice-prune-gate). Durable outcomes are already graduated: the sweep itself, the upstream-drift CHANGELOG 1.6.2 record, and follow-up issues #3525 and #3526; the PR body carries the contract summary. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_015DrfSCgSk56RoVv2gKWmQm
|
Claude finished @kyle-sexton's task in 2m 49s —— View job Code Review (
|
|
Claude finished @kyle-sexton's task in 1m 15s —— View job Security review —
|
…-out 5 A repo-wide /provenance:audit over 1292 tracked markdown files, plus the two plugin changes it justified and the documentation remediation it produced. Plugin behavior (provenance 0.4.2 -> 0.5.0): - Enable trigger_less_stamp_check in .claude/provenance.json. The portable default is off, so a dated stamp carrying no observable recheck trigger had never been flagged in this repo. Enabling it surfaced 152 such stamps, 30% of the 500 that parse here. A date is never authority: a stale stamp reads identically to a fresh one, and the trigger is the load-bearing part. Remediation is per-stamp judgment, so the rule reports rather than auto-applies. - Widen excluded_paths to **/evals/fixtures/**. Planted copies in a fixture tree are the harness working, not a provenance defect; the previous glob left 60 such files audited as ordinary prose, one of which reached a split judge panel. - Rubric version 4 narrows carve-out 5 with a span-level qualifier. Version 3 asked only whether a surface's product is a distillation, which nearly any reference file crediting a source can answer yes to; because carve-outs are graded before the criteria and stop grading, the broad reading absorbed C2 and C4 failures the criteria exist to catch. Version 4 keeps the purpose test and adds the boundary a distilling file draws for itself. This invalidates the version-3 golden-set measurement recorded in 0.4.0; the set must be re-scored before any precision figure is cited. Documentation remediation (17 surfaces, all additive): - 15 dated surfaces given observable recheck triggers, each naming events specific to what that file depends on. Findings 152 -> 91. - docs-hygiene/skills/audit-noise/SKILL.md: attribute the negation rule's worked example to Anthropic's Claude 4 best-practices page, which the file's Sources section did not list. Found by a unanimous three-judge panel on C3. - kindle-dedrm/skills/manage/references/sources.md: record that the primary tutorial URL and the successor a previous migration moved it to are both dead, using the demotion the plugin's own dispositions reference prescribes. Eight plugins whose shipped content this change set modifies are bumped with per-plugin changelog entries, because the directory-source install cache keys on manifest semver rather than commit SHA: without the bump every remediation here would be undeliverable to consumers already on the current version. Verification: scripts/affected-tests.sh --run passes 126 shell suites; the 4 selected Python suites pass 458 tests; fingerprint.test.mjs 40/40; markdownlint, typos and editorconfig clean. Every documentation change is purely additive. A blind semantic-diff guard caught two defects excluded from this change: a trigger insertion that split a verdict paragraph, and an edit written to satisfy the detector's regex rather than inform a reader. Refs #3525, #3524, #2297 Co-Authored-By: Claude <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_017pyWQJR7BoENh1131RPy2Y
Resolves two conflicts, both mechanical: plugin.json: keep 0.39.0. This branch's minor bump supersedes main's 0.38.22 patch; the two changes are independent and both ship. CHANGELOG.md: keep both sections, 0.39.0 above 0.38.22. Neither entry replaces the other. Suite 93/93 on the merged tree. NOTE, unresolved by this merge and tracked separately: main's #3524 added a SKILL.md paragraph and a plugin.json description clause restating the least-invoked-first drop order and citing the official skills page for it. This branch's own description now says decay-weighted, so the merged file states both. An adjudicator is settling the wording; the contradiction is deliberate and visible rather than silently resolved one way here. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_015eyw6KUwExd78yyptowV6d
…, and date it The correction stays at full strength, but "stated wrongly" with no referent put the error on the ADR's author, who transcribed the vendor page near-verbatim. It also dropped the operationally useful fact: the upstream page is still wrong as of 2026-08-31, which is why the same sentence keeps re-entering this repo. It is on main in the skill's SKILL.md and in the plugin manifest, and #3524 added a citation pointing at the wrong page for it. The SKILL.md paragraph now separates the two upstream pages by whether they hold: the settings page owns the budget fraction and per-entry cap and matches; the skills page states the drop order and does not. Readers are routed to the binary for the ordering. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_015eyw6KUwExd78yyptowV6d
…-out 5 A repo-wide /provenance:audit over 1292 tracked markdown files, plus the two plugin changes it justified and the documentation remediation it produced. Plugin behavior (provenance 0.4.2 -> 0.5.0): - Enable trigger_less_stamp_check in .claude/provenance.json. The portable default is off, so a dated stamp carrying no observable recheck trigger had never been flagged in this repo. Enabling it surfaced 152 such stamps, 30% of the 500 that parse here. A date is never authority: a stale stamp reads identically to a fresh one, and the trigger is the load-bearing part. Remediation is per-stamp judgment, so the rule reports rather than auto-applies. - Widen excluded_paths to **/evals/fixtures/**. Planted copies in a fixture tree are the harness working, not a provenance defect; the previous glob left 60 such files audited as ordinary prose, one of which reached a split judge panel. - Rubric version 4 narrows carve-out 5 with a span-level qualifier. Version 3 asked only whether a surface's product is a distillation, which nearly any reference file crediting a source can answer yes to; because carve-outs are graded before the criteria and stop grading, the broad reading absorbed C2 and C4 failures the criteria exist to catch. Version 4 keeps the purpose test and adds the boundary a distilling file draws for itself. This invalidates the version-3 golden-set measurement recorded in 0.4.0; the set must be re-scored before any precision figure is cited. Documentation remediation (17 surfaces, all additive): - 15 dated surfaces given observable recheck triggers, each naming events specific to what that file depends on. Findings 152 -> 91. - docs-hygiene/skills/audit-noise/SKILL.md: attribute the negation rule's worked example to Anthropic's Claude 4 best-practices page, which the file's Sources section did not list. Found by a unanimous three-judge panel on C3. - kindle-dedrm/skills/manage/references/sources.md: record that the primary tutorial URL and the successor a previous migration moved it to are both dead, using the demotion the plugin's own dispositions reference prescribes. Eight plugins whose shipped content this change set modifies are bumped with per-plugin changelog entries, because the directory-source install cache keys on manifest semver rather than commit SHA: without the bump every remediation here would be undeliverable to consumers already on the current version. Verification: scripts/affected-tests.sh --run passes 126 shell suites; the 4 selected Python suites pass 458 tests; fingerprint.test.mjs 40/40; markdownlint, typos and editorconfig clean. Every documentation change is purely additive. A blind semantic-diff guard caught two defects excluded from this change: a trigger insertion that split a verdict paragraph, and an edit written to satisfy the detector's regex rather than inform a reader. Refs #3525, #3524, #2297 Co-Authored-By: Claude <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_017pyWQJR7BoENh1131RPy2Y
…-out 5 (#3531) No linked issue ## Summary A repo-wide `/provenance:audit` run over 1352 tracked markdown files, plus the two plugin changes it justified. The copy-detection lane found almost nothing actionable; the finding that matters came from a rule that ships **off** by default. Enabling it surfaced 152 dated verification stamps carrying no observable recheck trigger, 30% of the 500 that parse in this corpus. This PR enables that rule, narrows a rubric carve-out the run showed was absorbing real findings, and gives 17 dated surfaces the triggers they were missing. ## Fix **Plugin behavior (provenance 0.4.2 → 0.5.0)** - **`trigger_less_stamp_check` enabled** in `.claude/provenance.json`. The portable default is off, so a dated stamp with no trigger had never been flagged here. This is the defect the upstream-drift convention names directly: a date is never authority, a stale stamp reads identically to a fresh one, and the trigger is the load-bearing part. Remediation is per-stamp judgment, so the rule reports rather than auto-applies. - **`excluded_paths` widened** from this plugin's own eval fixtures to `**/evals/fixtures/**`. Planted copies in a fixture tree are the harness working; the previous glob left 60 such files in other plugins audited as ordinary prose, and one reached a split judge panel. Corpus 1352 → 1292. - **Rubric version 4: carve-out 5 gains a span-level qualifier.** Version 3 asked only whether the surface's product is a distillation, which almost any reference file crediting a source can answer yes to. Carve-outs are graded before the criteria and stop grading, so the broad reading absorbed C2 and C4 failures the criteria exist to catch. Version 4 keeps the purpose test and adds the boundary a distilling file draws for itself: a verbatim span its own attribution does not enumerate is not covered, and reformatting is not distillation. **This invalidates the version-3 golden-set measurement recorded in 0.4.0** (8 tp / 0 fp / 0 fn / 2 tn, precision 1.00); the set must be re-scored against version 4 before any precision figure is cited, per the rubric's own versioning rule. Evidence lives in `CHANGELOG.md`, never in the rubric, which is inlined into every judge prompt. **Documentation remediation (17 surfaces, all additive)** - **15 dated surfaces given observable recheck triggers**, each naming events specific to what that file depends on: a settings key changing name or default, a cited page changing the section a claim rests on, a Suno release note revising a character-limit figure, a SARIF spec revision renumbering the sections a finding identity keys to. Findings 152 → 91. - **`docs-hygiene/skills/audit-noise/SKILL.md`**: the negation rule's worked example comes from Anthropic's Claude 4 best-practices page, which the file's Sources section did not list. Added. Found by a unanimous three-judge panel on C3. - **`kindle-dedrm/.../references/sources.md`**: the primary tutorial URL this skill depends on 404s, as does the successor a previous migration moved it to. Recorded with the demotion the plugin's own dispositions reference prescribes: keep the claims, mark the basis dead, name a recovery trigger. No archived snapshot exists. ## Verification Re-run after merging `main` (which landed #3524 and provenance 0.4.2): - `scripts/affected-tests.sh --run`: **126 shell suites pass**, 0 fail. - The 4 selected Python suites in their own lane: **458 tests pass** (`test_hygiene.py` 317, `test_observer.py` 72, `test_prune_babysit_worktrees.py` 45, `test_check_contract_clause_coverage.py` 24). `fingerprint.test.mjs`: 40/40. - `markdownlint-cli2` on all changed markdown: 0 issues. `typos`, `editorconfig-checker`, and `check-purged-em-dashes.sh` all clean. - Every touched surface re-checked with `check-stamps.sh`: findings 0, `counts.parsed` unchanged, so a trigger was added and no stamp removed. - Every documentation change is **purely additive** (`git diff --numstat` shows no deleted content lines): no claim, date, or existing sentence altered. - A blind semantic-diff guard ran over the batch and caught two defects that are **not** in this PR: a trigger insertion that split a verdict paragraph (repaired, `12948c25`), and an edit whose stated purpose was satisfying the detector's regex rather than informing a reader (reverted). - **Composes with #3524 rather than colliding**: no file overlap, and its four-part stamp upgrades raise parsed stamps 501 → 525 while clearing 5 of this PR's findings (96 → 91). The only merge conflicts were the shared version bump and adjacent CHANGELOG entries, both resolved additively. ## Related - Refs #3525 — provenance redesign. Two comments there carry this run's evidence: the recall lane already nominated 17 of the 18 surfaces #3524 remediated (so the gap is downstream routing, not the nominator), and a per-lane precision table arguing where the script/model line belongs. - Refs #3524 — the frontmatter-alignment sweep, now merged. Complementary: it cleared the frontmatter-fact class, this clears part of the trigger-less class. - Refs #2297 — tracked unstamped-carrier intake. This PR clears part of the trigger-less class, not all of it: **91 findings remain**, and roughly 20% of those are detector false positives (files stating a real trigger in prose the regex cannot match, e.g. "enqueue when harness support lands"). Two such files are deliberately left flagged rather than papered over, so the detector gap stays visible. The fix is to demote that check from a judge to a nominator, argued in #3525. - `docs/conventions/upstream-drift/README.md` — the convention this enforces. 🤖 Generated with [Claude Code](https://claude.com/claude-code) https://claude.ai/code/session_017pyWQJR7BoENh1131RPy2Y Co-authored-by: Claude <noreply@anthropic.com>
…bric and relay row (#5439) Closes #3525 ## Summary The attribution audit could only act on lexical copies, so a restated constant (a default, a limit) in a paraphrase never reached a panel or the relay. Per the owner decision on #3525 (Option 1), this keeps the nominator and adds a report-only restated-fact lane beside the copy lane. The `--diff` quality-gate mode is not part of this change. ## Fix - Second rubric, restated-fact v1, in `reference/rubric.md`; the copy rubric stays at v4. The dispatching run routes a candidate, not the nominator's class guess, and an unnamed dispatch defaults to copy. - Every restated-fact `STANDS` goes through one refutation pass; a refuted or open finding stays on the human report. Attack (e) treats only a pointer that stands in place of the fact, or a whole four-part record, as conforming; a link beside a stated value is not one. - New finding class `restated-fact`, rule id `attribution/audit/rule-restated-upstream-fact`, never fix-eligible, with report-only dispositions. - Relay row for a unanimous, upheld, refutation-surviving restated fact; ranks last, tier IMPORTANT, empty confidence. Only the fully qualified rule id relays; a record carrying the slug under another prefix is withheld and counted, as the copy and stamp rules match by slug, and never reaches `## Unparsed`. - Golden cases c11 to c24 (six positives, eight negatives including stamped records, a conforming pointer and distilling files) and evals 11 to 13. The six positives are a subset of the 11 bare restatements #3524 converted: the rest fail the restated-fact rubric or an existing carve-out (history-only sentences, owned measurement content, quoted-and-cited spans), and the `plugin-quality` command is used as the post-sweep pointer negative (c19). - `attribution` 0.8.0 and `review` 0.34.2 with CHANGELOG entries; detector-findings convention 3.5.0. ## Verification - `node plugins/attribution/skills/audit/scripts/fingerprint.test.mjs`: 40 passed, 0 failed. - `bash scripts/affected-tests.sh --run`: exits 1 only on `scripts/check-script-contract.test.sh`, whose two `check-html-assets.sh` cases need `node_modules/.bin/htmlhint` (not installed in this worktree; CI installs it). Every other suite it runs passes. It also lists the fingerprint suite among the five it does not run, which was run separately above. - After merging `origin/main` (attribution 0.8.0 and review 0.34.2, above main's 0.7.1 and 0.34.1): `bash scripts/check-changelog-parity.sh --check-bump origin/main` and `--check --check-order`, `bash scripts/check-stale-base-overlap.sh --check origin/main`, `bash scripts/validate-plugins.sh` and `node scripts/validate-plugin-contracts.mjs` pass. - `npx markdownlint-cli2` over every markdown file changed on the branch: 0 issues. - `typos` over every file changed on the branch: clean (the lint job failed earlier on `neighbouring`, now `neighboring`). - `emit-findings.test.sh` 562/0, `score-golden.test.sh` 41/0, `check-detector-findings-crosswalk.sh --check`, `check-detector-eval-coverage.sh`, `check-fixture-git-isolation.sh`, and `check-skill.sh` (0 errors, no new warning) pass. ## Related - #3524 (the census the golden cases are drawn from). - `plugins/review/skills/audit-enforceability/context/crosswalk.md` gained a rule-id row for the new rule; the review plugin is bumped to 0.34.2 for it. 🤖 Generated with [Claude Code](https://claude.com/claude-code) --------- Co-authored-by: Claude Sonnet 5.5 <noreply@anthropic.com>
No linked issue
Summary
Frontmatter-alignment sweep: every living surface that states a skill-frontmatter fact now either points at the official frontmatter reference so readers fetch the latest, or carries a conforming four-part record (claim, basis, as-of date, observable recheck trigger) where restating is necessary to function; the skill fleet's invocation posture is now uniformly explicit. Grounded in a verbatim rung-1 fetch of the current skills page (2026-08-31) and a verified repo census (11 bare prose restatements, 3 undated constant-encoding scripts).
Fix
check-skill.sh,check-listing-budget.sh, andaudit_skill_visibility.pycarry dated, triggered records beside their encoded constants (scripts cannot fetch upstream at runtime), modeled on the conforminglevers.jsonexemplar.namereplaced the whole command" history is re-anchored to the changelog basis in PLUGIN-PHILOSOPHY and the skill-authoring playbook, since the skills page no longer pins it (it now records the v2.1.216–v2.1.245 prefix-doubling quirk fixed in v2.1.246).user-invocable: trueadded to the 5 skills missing it (documented default, zero behavior change, matches the explicitdisable-model-invocationconvention).Verification
scripts/affected-tests.sh --run: 140 shell suites pass pre-merge and 139 pass re-run after integrating origin/main, 0 fail.check-skill.shpasses (rc=0, 0 FAILs) on all 8 touched skills;check-listing-budget.shstill runs clean after its comment changes.1,536|1536|skillListingBudgetFraction|skillListingMaxDescCharsover docs/plugins/scripts minus recorded exclusions): zero unstamped carriers remain.stale-base-overlap-gate(origin/main merged in; the overlapping claude-ops CHANGELOG change was a link fix in an old entry, merged cleanly) andcontract-slice-prune-gate(slice pruned).Related
docs/conventions/upstream-drift/README.md(the convention this sweep applies) and its 1.6.2 CHANGELOG entry recording the fired trigger.paths/near-cap/when_to_usecandidates).🤖 Generated with Claude Code
https://claude.ai/code/session_015DrfSCgSk56RoVv2gKWmQm