fix(skills): give every skill description real 'Use when:' trigger phrasing - #2181
Conversation
|
Bugbot is not enabled for your account, so this pull request was not reviewed. Enable Bugbot in the Cursor dashboard to get automatic reviews on future PRs. |
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: 7b9376a488
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
7b9376a to
887be02
Compare
|
Claude finished @kyle-sexton's task in 4m 29s —— View job Review complete
Findings (posted inline)
What I independently checked (not just author-claimed)
Not independently verified (author-claimed only)
No other correctness, security-lane, or maintainability issues found in the changed files — this is a low-risk, mechanical description/changelog/version-bump sweep and the two items above are the only substantive findings. Branch: |
|
Claude finished @kyle-sexton's task in 1m 49s —— View job Security review — no issues found
Scope reviewed: the full PR diff ( Findings: none. Every change in this PR is confined to YAML frontmatter
On the Codex bot comment (kindle-dedrm |
|
Claude has reviewed this PR 2 times. The lane skips further automatic reviews after 5; deleting this comment resets the count. |
`plugins/verification/skills/measure/SKILL.md`'s frontmatter `description` opened with what the skill IS and buried its routing phrases behind a lowercase `use for`. A description is the text Claude matches against to decide whether to LOAD a skill, so a summary-shaped one makes the skill under-fire: the skill-quality gate reported "description has no 'Use when:' trigger phrasing — a description is a trigger spec, not a summary". Move the phrases behind `Use when:` in the marketplace's house shape and add six more a user would actually type. Every single-quoted phrase the previous description carried is kept verbatim — including `'cannot quantify'`, which is prose the trigger extractor also tracks — so the gate's keyword- preservation check sees a superset rather than a rewrite. Refs #2174 Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01FdBsorqGjCU4Mg5nrmgz4m
…sing
A skill `description` is the text Claude matches against to decide whether to
LOAD the skill, so a summary-shaped description makes the skill under-fire.
The skill-quality gate has two checks for this, and a corpus-wide sweep found
28 skills tripping one of them beyond `verification:measure`:
* 20 more carried no `Use when:` phrasing at all — most had good routing
phrases sitting behind a lowercase `use for`, which the gate does not
recognize; a few named only topics or states, never a phrase a user types.
* 7 had `Use when:` but no SINGLE-quoted phrases. The gate's trigger-drop
protection tracks only `'single-quoted'` text, so phrases written with
escaped double quotes (debugging, docs-hygiene, kindle-dedrm) or as a bare
comma list (repo-hygiene) had no regression cover at all — a later rewrite
could silently drop any of them.
Fix 25 of the 28, across context7, debugging, docs-hygiene, firecrawl,
implementation, kindle-dedrm, planning, playbooks, repo-hygiene, review,
testing, toolchain and verification. Every phrase already present is preserved
verbatim, so the gate's keyword-preservation check sees a superset rather than
a rewrite; each list also gains phrases a user would actually type rather than
a paraphrase of the skill's purpose.
Four are deliberately left as-is, all `disable-model-invocation: true` with no
pre-existing quoted phrase: discipline:wait-what, firecrawl:update,
github:setup, playbooks:update. A skill that cannot be model-invoked is never
matched against its description, so triggers would buy no auto-invocation
there — and two of them are maintainer-only drift checks that say "Not for
consumers", where user-typed triggers would invite exactly the misfire this
issue is about. Where a dmi:true skill already had quoted phrases
(planning:questionnaire), the connector is normalized without inventing any.
Closes #2174
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01FdBsorqGjCU4Mg5nrmgz4m
887be02 to
b807761
Compare
|
Claude finished @kyle-sexton's task in 3m 1s —— View job Review complete
What changed since the last review roundThe diff at the current head (
New findings (this round)
Not independently verified (author-claimed only)Per lane instructions, these are not restated as confirmed: the "192/196 skills clean" full-corpus sweep, the "0 collisions" cross-skill trigger check, the "109 phrases added / 0 dropped" count, and the CI gate results table. I did verify by direct diff inspection that no existing single-quoted trigger phrase was removed anywhere in the 25 changed No correctness, security, or maintainability issues found beyond the three items above — this remains a low-risk, mechanical frontmatter/changelog/version-bump sweep. Branch: |
|
Claude finished @kyle-sexton's task in 1m 1s —— View job Security review — no issues found
Scope reviewed: the full PR diff at HEAD Findings: none. Every hunk in this diff is confined to:
None of that constitutes an injection surface, a permission/trust-boundary change, or a place secrets/tokens could leak:
On the two functional threads already raised on this PR (Codex's and my own prior code-review comment, re: |
…eferred gate calls (#2186) No linked issue ## Summary Three deferred judgment calls, all made here. **The uncited doc pages were never dispositioned, and the set is larger than recorded.** PR #2177 worked from an index of 112 core `/docs/en/` pages with ~33 uncited. Re-derived today, the index lists **135** core pages and this repo cites **78**, leaving **57** unevaluated. #2177 also demonstrated the miss rate of dismissing by eye: it took six pages that had been written off as irrelevant and found every one load-bearing enough to warrant a verdict. So all 57 get a recorded disposition here rather than a second eyeball pass. The finding that mattered is in the **cross-platform contract**. It reads one axis — the operating system — and names `feature-availability` as its canonical input. That page carries two axes, model provider and subscription plan, and scopes itself to what runs locally: "The Claude Code CLI and everything that runs locally work on every provider." The **host surface** a consumer runs in was therefore never read at all, by either the contract or its input. It has to be, because a host can withhold the plugin system itself rather than one capability, and where no plugin loads there is no portable path for one to owe. **`check-skill.sh` check 5** and **check 12's 4-skill warning floor** were both left open as "a separate call". Both are decided, at their own sites, with the reasoning recorded so neither is re-litigated from a false premise. ## Fix ### 1. Uncited-page disposition (57 pages) One stated relevance test, applied to all 57 so the dismissals are auditable rather than tacit: > **Relevant** if the page describes a surface a plugin author can **declare, invoke, or must > accommodate.** Otherwise **not relevant.** **Split: 1 adopt / 0 defer / 3 decline / 4 relevant-as-evidence / 49 not relevant.** Verdicts land in `docs/PLUGIN-PHILOSOPHY.md` under Native-first → **Recorded gate runs**, in the [upstream-drift](https://github.com/melodic-software/claude-code-plugins/blob/main/docs/conventions/upstream-drift/README.md) four-part shape (claim, basis, as-of date, recheck trigger), following #2177's form. | Page | Disposition | |---|---| | `platforms` | **Relevant → ADOPT, as a citation.** The canonical host-surface index; the axis `feature-availability` does not carry. Lands doctrine. | | `github-enterprise-server` | **Relevant → DECLINE.** A real plugin-distribution surface ("Plugin marketplaces \| ✅ Supported") — declines on need, not subject: nothing here documents a GHES-hosted mirror or fork. | | `ultrareview` | **Relevant → DECLINE.** Fails gate 1: every run is human-confirmed and metered, so no skill can reach it. | | `chrome` | **Relevant → DECLINE.** Ships as the built-in `claude-in-chrome` skill; nothing to declare. Recorded because it only *looked* cited — see the dead-link note below. | | `desktop`, `vs-code`, `mobile`, `desktop-wsl` | **Relevant → read in full as the evidence base for the `platforms` row**; no separate verdict, because they are one finding seen from four pages rather than four surfaces. Quoted verbatim in that row. | | `jetbrains` | Not relevant — false friend: "plugin" there is the JetBrains IDE plugin, a different sense. Ranked 2nd by `plugin`-keyword density and is the one page where that signal is pure noise. | | `desktop-quickstart`, `desktop-linux`, `desktop-ios-simulator`, `web-quickstart`, `troubleshoot-install` | Not relevant — install and first-run recipes; nothing declarable. | | `slack`, `claude-tag` | Not relevant on their own — delegation front ends indexed by `platforms`, which is the adopted citation; `slack` is additionally being retired for Team/Enterprise. | | `devcontainer` | Not relevant — container recipe; its only `marketplace` hit is a VS Code extension link. | | `gitlab-ci-cd`, `github-actions-cloud-providers` | Not relevant — CI recipes and provider IAM routing; zero plugin or skill surface (`gitlab-ci-cd`: 0 keyword hits). | | `amazon-bedrock`, `google-vertex-ai`, `microsoft-foundry`, `claude-platform-on-aws` | Not relevant — provider auth/IAM config; the plugin-facing consequence is the availability matrix, already adopted as `feature-availability`. | | `gateways`, `llm-gateway`, `llm-gateway-connect`, `llm-gateway-protocol`, `llm-gateway-rollout` | Not relevant — org request-routing plane between the client and a provider; no plugin declares or observes it. | | `claude-apps-gateway`, `claude-apps-gateway-config`, `claude-apps-gateway-deploy`, `claude-apps-gateway-on-aws`, `claude-apps-gateway-on-gcp`, `claude-apps-gateway-spend-limits` | Not relevant — deploying and operating Anthropic's gateway product; `gateway.yaml`, Kubernetes, spend caps. | | `self-hosted-environments`, `self-hosted-environments-quickstart`, `self-hosted-environments-configuration`, `self-hosted-environments-deploy`, `self-hosted-environments-identity`, `self-hosted-environments-reference`, `self-hosted-environments-testing` | Not relevant — standing up and operating cloud-session runners on org infrastructure. | | `admin-setup`, `authentication`, `legal-and-compliance`, `third-party-integrations` | Not relevant — enterprise deployment, identity, and policy plane; no surface a plugin declares or observes. | | `analytics` | Not relevant — but **fetched, not assumed**, because per-skill or per-plugin cost attribution would have bound instruction economy. It has none: attribution is PR-level only. (The per-skill/per-plugin usage breakdown is a consumer-side `/usage` dialog, not an authoring input.) | | `network-config` | Not relevant, and the third clause of the test is why rather than the family label: proxy, custom CA, and mTLS are **transport configured on the client**, so a skill making a network call either succeeds or sees an ordinary failure — there is nothing to declare or degrade. Its two plugin-adjacent lines are egress allowlist entries a network admin sets, not a plugin (`downloads.claude.ai` for "Plugin executable downloads"; `storage.googleapis.com` for "plugin metadata shown in `/plugin`"). | | `corporate-launcher` | Not relevant, checked against the page rather than dismissed as admin tooling: `CLAUDE_CODE_PROCESS_WRAPPER` wraps "every process Claude Code launches **from its own binary** — the background service, every session it hosts in agent view, and Claude Code's relaunches after an update". A plugin's `${CLAUDE_PLUGIN_ROOT}/bin/` invocation is a Bash-tool subprocess, not a Claude Code self-spawn, so the `bin/` stance is unaffected and owes no change. | | `champion-kit`, `communications-kit` | Not relevant — internal-advocacy and rollout-comms collateral. | | `accessibility`, `keybindings`, `terminal-config`, `voice-dictation`, `fullscreen`, `fast-mode` | Not relevant — consumer client settings; no plugin declares or must accommodate them. | | `prompt-library` | Not relevant — copy-paste prompts for users, not an authoring surface. | **Doctrine added — one paragraph, plus four table rows.** The cross-platform contract gains the host axis, citing `platforms` and restating none of its facts. The three verbatim host facts (Desktop-in-WSL sessions lack "connectors and plugins"; `/plugin` "[doesn't] work from the app" on mobile; Desktop's Cowork tab sources plugins "not from the CLI's `~/.claude` directory") live in the gate-run row, where they carry a recheck trigger — not in the contract, which states only the rule they establish. **A dead citation, deliberately not fixed.** Every doc URL this repo cites was checked live — all 81 slugs plus the 4 subpath citations (`agent-sdk/overview`, `agent-sdk/agent-loop`, `agent-sdk/plugins`, `whats-new/2026-w32`). **84 of 85 return 200.** One does not: `code.claude.com/docs/en/browser` now 404s (`chrome` is the live page). Its sole occurrence is `plugins/playbooks/skills/boris/vendor/SKILL.md:938` — a **verbatim upstream baseline kept for drift detection**, which the plugin README says to treat as untrusted and which `/playbooks:update` owns. Hand-editing it would corrupt the vendor SHA it exists to compare. Recorded in the `chrome` row with that path as its recheck trigger instead. ### 2. `check-skill.sh` check 5 — KEEP the extractor as-is (decided, recorded at the site) Two premises are usually offered for narrowing to markdown-link targets. Both are false, and the comment now says so, because the premise is what keeps the question alive: 1. **"It matches bare paths in prose."** It does not, and never did. Both generators are delimited — backtick-wrapped, or a `](…)` link target — and both are scoped to the `INTERNAL_DIRS` allowlist. Naked prose cannot match. (#2179's own summary and CHANGELOG entry describe it as extracting "prose and inline-code refs"; the in-script wording is corrected here to match what the greps do.) 2. **"The backtick branch is redundant."** Measured over the 196-skill corpus rather than argued: | Measure | Count | |---|---| | Backtick-form refs, all SKILL.md | 282 | | Link-form refs, all SKILL.md | 475 | | **Unique backtick-form refs with no link form anywhere in the same file** | **122** | | …spread across | **39 skills** | | …of those 122, resolving to a real file today | **122 (100%)** | Narrowing would drop 122 real, currently-resolving supporting-file references across 39 skills. The link branch being the larger share is not the question; the overlap is, and 122 refs sit outside it. The false-positive risk that motivated the proposal is real but **latent, not observed** — zero on the current corpus. It is handled by message wording (every failure carries `hand-verify the line before fixing, may be an illustrative example`) rather than by deleting coverage of 39 skills. Reopen only if a false positive is actually observed. ### 3. Check 12's 4-skill warning floor — INTENTIONAL, no dmi carve-out (all 4 confirmed) #2181's reasoning holds, and upstream states the premise more strongly than #2181 did. The skills doc's frontmatter-behavior table gives, for `disable-model-invocation: true`: **"Description not in context, full skill loads when you invoke"** — so trigger phrasing on such a skill cannot route anything, at all. `user-invocable` defaults to `true` (confirmed on the same page, not assumed), so `github:setup` omitting it is slash-command-only, exactly its declared contract. The load-bearing half of #2181's argument is the *stranded-phrase* test, which is an empirical claim about the current tree, so each was re-checked against the tree rather than against #2181's prose: | Skill | Verdict | Confirmed against the tree | |---|---|---| | `discipline:wait-what` | **Right to leave** | Its description *is* the instruction; the trigger is noticing you have stopped following. No sibling needed — by construction the model cannot detect it. | | `firecrawl:update` | **Right to leave** | Maintainer-only. Sibling `firecrawl:firecrawl` **verified** to carry the consumer phrases (`'scrape this page'`, `'crawl this site'`, `'WebFetch is blocked'`, …). Nothing stranded. | | `playbooks:update` | **Right to leave** | Maintainer-only. Sibling `playbooks:boris` **verified** to carry `'how does Boris use Claude Code'`, `'Claude Code workflow tips'`, `'optimize my CLAUDE.md'`, … Nothing stranded. | | `github:setup` | **Right to leave** — the weakest of the four as originally argued, and it holds | #2181 argued from intent ("user-invoked only"). Checked instead for a stranded phrase: model-invocable siblings `github:advise` and `github:audit` carry the plugin's consumer-facing routing, including `'help me set up Y'`. `setup` covers plugin *prerequisites* (gh auth, writing `.claude/github/`), which is a deliberate slash command, not a routing target. | **No carve-out is added**, and that is the recorded call. Exempting dmi-true from check 12 would suppress a warning that is doing no harm while hiding the `kindle-dedrm` failure mode #2181 itself surfaced — a phrase reachable only from a skill the model can never match. The floor stays; the exemptions stay documented at the check-12 site. ## Verification **Method.** Every page was fetched with `curl -sL …/<slug>.md` — the raw markdown, not WebFetch. That removes the summarizer and the truncation window from the loop entirely, so the METHOD RULE holds trivially: every upstream sentence quoted in this PR and in the doctrine is verbatim from a complete page, and a genuine "the page never states X" is a checkable claim rather than a routine false negative. Byte counts confirm no truncation (e.g. `desktop.md` 96,288 bytes, `vs-code.md` 49,764). No page was asked to confirm a sentence from this repo. **The uncited set was re-derived, not inherited.** The grep was also re-run with **no `--include` filters** to be sure no citation lives in a file type the filter misses — identical result, 81 slugs, so 57 uncited is the real number. **Every cited URL was checked live**: 84 of 85 (81 slugs + 4 subpath citations) return 200; the single 404 is the vendored `browser` link described above. **The gate-1 check that decided the headline adopt** was run against the page rather than assumed: `feature-availability`'s section headings are *Availability by model provider*, *Availability by subscription plan*, and *Model availability* — no host-surface axis — and its only feature table header row is `| Feature | Pro | Max | Team | Enterprise |`. Had it carried a host axis, `platforms` would have been a redundant second index and this would be a decline instead. Gates run the CI way, against the **committed** tree, base-ref form: | Gate | Result | |---|---| | `bash scripts/check-contract-slice-prune.sh --check-diff origin/main` | pass — leaves no path under `docs/topics/` | | `bash scripts/check-changelog-parity.sh --check-bump origin/main` | pass | | `bash scripts/check-changed-skills.sh origin/main` | pass — no changed skills | | `bash scripts/check-skill-portability.sh origin/main` | pass — no skill files in scope | | `bash scripts/check-shell-portability.sh origin/main` | pass — no unexcused GNU-only constructs | | `npx --yes markdownlint-cli2` over all 3 changed `.md` | **0 errors** | | `shellcheck` + `shfmt -i 2 -d` on `check-skill.sh` | clean | | `bash -n check-skill.sh` | clean | | Line endings | all 5 changed files `i/lf w/lf` | Because the change to `check-skill.sh` is comments only, `check-changed-skills.sh` exercises nothing — so the script was run directly to prove it still parses and behaves: - `check-skill.sh measure` → `PASS — 0 errors, 0 warning(s)`, `all 10 base-ref trigger phrase(s) preserved`. - `check-skill.sh wait-what` → `PASS — 0 errors, 2 warning(s)`, one of which is verbatim `description has no 'Use when:' trigger phrasing` — confirming the documented floor still fires as described rather than being silently suppressed. - `check-skill.test.sh` runs to completion in CI (`plugin-gate`); locally on Windows/Git Bash it is impractically slow, per the coverage note #2179 recorded. Nothing here is behavioral. `plugins/skill-quality` → **0.15.2** with a matching `## [0.15.2]` entry. The `docs/` changes are docs-only and owe no plugin bump; the `upstream-drift` **Adopters** registry already carries a row for the gate-run table (added in #2177), and these rows join that table rather than create a new adopter, so that convention needs no version change. `docs/OFFICIAL-DOCS.md` gains the four newly load-bearing pages, per the rule its own warning states and the precedent #2177's review set: a needed page that is not listed must be added. No `docs/topics/<slug>/` directory was created — the durable outcome is doctrine text, as the Contract-tier prune rule requires. ## Review rounds Three threads, all real, all answered and resolved. Each found a defect in the *basis* of a row rather than in its verdict, which is the failure mode a decision record most needs caught: a verdict outlives the reasoning nobody re-reads. - **The GHES row's premise was overstated and its trigger fired on arrival** (`chatgpt-codex-connector`). It claimed "every plugin README ships the github.com shorthand". Re-derived from the tree: 54 of 65 carry the literal string, 9 carry no install block, `dometrain` points at another github.com marketplace, and `plugins/github/README.md` — deliberately marketplace-agnostic — uses the `<marketplace-owner>/<marketplace-repo>` placeholder. All are still `owner/repo` shorthand, so the trigger now names the form that actually signals a non-github.com host, a **full git URL**, of which the tree has none. The verdict stays Decline, but the review surfaced a real finding that had been waved through and is now recorded in the row: a consumer redistributing the `github` plugin from a GHES-hosted marketplace would follow that README and have the shorthand silently resolve to github.com instead of their own instance. - **"Platform" was doing two jobs** (`chatgpt-codex-connector`). The existing `feature-availability` row and `docs/OFFICIAL-DOCS.md` both described that page as covering "platform, provider, and plan", while this change rests on the host axis being absent from it. Both senses of the word in one table would let a future audit read the host axis as already covered and retire the new row as redundant. The page's own sense is the **provider** platform — its axis headings are *Availability by model provider* and *Availability by subscription plan* — and both sites now say so explicitly. - **The `platforms` row claimed four evidence pages and quoted three** (`claude`). Correctly diagnosed as a missing fact rather than an overstated page: `vs-code` does carry a host-axis fact, and the most directly plugin-relevant of the four — its CLI-vs-extension table gives `Commands and skills` as `All` for the CLI against `Subset (type / to see available)` for the extension, so a skill this fleet ships may not be reachable there. It is now quoted in the row. All gates and `markdownlint-cli2` re-run clean over the changed files after these edits. CI is green, including `plugin-gate` — which runs `check-skill.test.sh`, the only executable proof that the check-5 comment insertions changed no behavior. ## Related - #2177 — established the Recorded-gate-runs table and its four-part row form; this run extends it and corrects its page census (112 → 135 core pages, ~33 → 57 uncited) - #2179 — deferred the check-5 extractor question as "a separate call"; decided here - #2181 — swept 196 skills for `Use when:` phrasing and left 4 with stated reasoning; all 4 re-reviewed and confirmed here - #2169 — the gate doc-currency audit these findings trace back to --------- Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
Closes #2174
Summary
A skill
descriptionis the text Claude matches against to decide whether to load the skill, soa description shaped like a summary of what the skill is makes the skill under-fire. The
skill-quality gate reported this on
verification:measure:This fixes
measureand sweeps the rest of the corpus, which the issue explicitly asks for. All196 skills were checked — not a base-ref diff. 29 carried a check-12 warning; 25 are fixed here and
4 are deliberately left as-is with the reasoning below.
Fix
Check 12 in
plugins/skill-quality/scripts/check-skill.shhas two branches, and the sweep found bothin the wild:
description has no 'Use when:' trigger phrasing'Use when:' triggers are not single-quotedNo
Use when:phrasing (22). Most already had good routing phrases, but behind a lowercaseuse forthe gate does not recognize (planning:*,testing:*,toolchain:check,review:fanout,implementation:implement,verification:confirm,verification:measure). A fewnamed only a topic or a state and no phrase a user types (
context7:lookup,toolchain:lint).Unquoted triggers (7). The gate's trigger-drop protection tracks only
'single-quoted'text, sophrases written with escaped double quotes (
debugging:debug,docs-hygiene:compress,kindle-dedrm:manage) or as a bare comma list (repo-hygiene:clean) had no regression cover atall — a later rewrite could have dropped any of them silently. Those are quoting changes; the
wording is unchanged.
Across the 25, 108 typed trigger phrases were added and zero existing phrases were dropped — the
gate's keyword-preservation check sees a superset, not a rewrite. This matters most on
measure,whose base-ref trigger set includes
'cannot quantify': that is prose, not a trigger, but theextractor tracks any single-quoted span, so the natural "clean this up" rewrite would have dropped it
and failed check 3. It is preserved verbatim.
Deliberately left as-is (4)
All four are
disable-model-invocation: trueand carried no pre-existing quoted phrase. A skillthat cannot be model-invoked is never matched against its description, so triggers would buy no
auto-invocation there — and inventing user-typed phrases for them would invite exactly the misfire
this issue is about.
The review round below sharpened this: dmi-true alone is not sufficient grounds, because a phrase
that lives only on a dmi-true skill is unreachable (that is the
kindle-dedrmfinding). So foreach of the four, the additional question is whether some phrase a user would type is left with no
model-invocable home. It isn't:
discipline:wait-what— its description is the instruction ("Type/discipline:wait-whatthemoment you notice you are skimming; only you know when you stopped following"). Self-observation
is the trigger; by construction the model cannot detect it.
firecrawl:update,playbooks:update— maintainer-only drift checks whose descriptions say "Notfor consumers — consumers update via
/plugin marketplace update". There is no user phrase thatshould route here, so nothing is stranded. Their model-invocable siblings (
firecrawl:firecrawl,playbooks:boris) are both fixed in this PR and carry the consumer-facing phrases.github:setup—SKILL.mdstates "User-invoked only", the plugin README lists it as"user-invoked only", and the only references to it are documentation. It is a deliberate slash
command, not an orphan and not a routing target.
Where a
disable-model-invocation: trueskill already had quoted phrases (planning:questionnaire),the connector is normalized to
Use when:and nothing is invented — it is the only dmi-true skillamong the 25 changed here.
One deliberate cross-skill duplicate
kindle-dedrm:managecarries a byte-identical copy of'set up Kindle DRM removal', a trigger onits sibling
kindle-dedrm:setup. It was invisible while double-quoted, and quoting it for trackingmakes two siblings claim the same typed phrase. It is kept anyway:
setupisdisable-model-invocation: true, so its description is never matched against user text, andmanage— model-invocable, with an action router that delegates to/kindle-dedrm:setup— is theonly skill that can receive the phrase by model invocation. Dropping the duplicate would leave the
phrase reachable only by an explicit slash command. A collision check over the whole corpus confirms
this is the only overlap: the other 107 added phrases collide with nothing.
Versioning
13 plugins touched, each with a patch bump and a matching
## [x.y.z]CHANGELOG entry:context70.5.1,debugging0.6.1,docs-hygiene0.10.1,firecrawl0.5.1,implementation0.13.1,
kindle-dedrm0.7.1,planning0.30.1,playbooks0.8.2,repo-hygiene0.9.1,review0.18.1,
testing0.5.2,toolchain0.12.1,verification0.5.2.Verification
How the sweep covered all 196 skills: the gate ships no
--allmode —check-changed-skills.shmapsa base-ref diff to skill dirs — so check 12 was reimplemented against the checker's own
plugins/skill-quality/scripts/skill-frontmatter.shsemantics (extract/field/strip_quotes/extract_triggers, including the intra-word-apostrophe normalization) and run over everyplugins/*/skills/*/SKILL.md. Two fidelity controls:check-skill.shon one skill from each warning class(
verification:measure,docs-hygiene:compress) — same classification both times.handling. No SKILL.md in the corpus uses a block-scalar
description:orwhen_to_use:(
grep -c '^description: *[|>]'over all 196 = 0), so both take the plain-scalar path.check-changed-skills.sh origin/mainruns the realchecker over exactly the 25 changed skills.
Gates run locally on the committed tree, base-ref form:
scripts/check-changelog-parity.sh --check-bump origin/mainscripts/check-changelog-parity.sh --checkscripts/check-skill-portability.sh origin/mainscripts/check-contract-slice-prune.sh --check-diff origin/mainmarkdownlint-cli2over all 38 changed.mdi/lf w/lfscripts/check-changed-skills.sh origin/main— the gate that reported the original warning — runsgreen in CI as
skill-quality-gateover all 25 changed skills. It was also run to completion locallyon
measurealone, before the sweep was layered on:PASS — 0 errors, 0 warning(s),all 4 base-ref trigger phrase(s) preserved.Post-change sweep: 192 of 196 skills clean, the remaining 4 being the documented exemptions.
Related
measure/context/metrics.md, not the frontmatter) and carried againas a noted follow-up on fix(session-flow,discipline): drop the per-session subagent cap that no longer exists, and record the doc-currency audit #2169
scripts/check-changed-skills.sh,plugins/skill-quality/scripts/check-skill.shcheck 12check-skill.shcheck 3