fix(skill-quality): reject unquoted description colons and overlong compatibility (Refs #3612) - #5045
Merged
Merged
Conversation
…ompatibility A plain description containing ": " does not parse as YAML, and Claude Code then loads the skill with no fields set. Check 1 now fails that form and fails a present compatibility value outside the spec's 1-500 characters. Quoted and block descriptions are unchanged, and an absent compatibility field stays valid. Co-authored-by: Kyle Sexton <kyle-sexton@users.noreply.github.com>
Contributor
|
PR body contract — issue linkage This PR body does not yet satisfy the issue-linkage contract:
Edit the body and this comment updates itself on the next run. |
The portability gate reads the escaped greater-than in that case arm as a GNU grep word boundary. Co-authored-by: Kyle Sexton <kyle-sexton@users.noreply.github.com>
This was referenced Sep 29, 2026
kyle-sexton
added a commit
that referenced
this pull request
Sep 29, 2026
…-invocation portability (#5307) Refs: #3526 Refs: #3612 Refs: #4586 ## Summary Fixes the skill-quality findings from the audit of the unattended Cursor agent's PRs (REPORT.md rows for #4586, #3612, #3526). `check` documentation and its own eval case now agree that check 25 is a blocking FAIL. Check 1 no longer hard-errors on provider-format `model` ids and now reads the whole `description` scalar. `measure-invocation score` fails clearly outside the marketplace checkout. The agent's unratified interop bullets are removed from the CHANGELOG. Ships as `skill-quality` 0.24.13. No `Closes`: #4586 and #3526 are already closed records, and #3612 keeps open acceptance (the operator's four interop answers). ## Fix - **Check 25 is a FAIL (#4586).** `evals.json` case 9 graded an advisory WARN, the opposite of `check-skill.sh`. README, the `check` step 3 output guide and the `check` description no longer call the polarity check advisory or say the skill never blocks. The manifest description names the `measure-invocation` harness. - **`check` gets `## Next`** and drops a paragraph that repeated the Arguments text (check 27b self-trip). - **Check 1, `model` (#4027, #3612).** Fails only an empty or whitespace-containing value; Bedrock ids with `:`, ARNs and Vertex ids with `@` pass. The skills page defines no stricter grammar. - **Check 1, `description` (#3612).** The unquoted-colon check covers every continuation line of a plain scalar and a line ending in `:`. Quoted and block scalars stay exempt. - **`measure-invocation` (#3526).** `score` exits 2 with guidance when no probe `skill_dir` resolves; `validate` adds an unresolved-count line; the `check` action resolves the probes directory explicitly. `reference/invocation-probes.md` states the lexical floor's limits (seed positives quote the description, no ingest path for plugin-eval or `claude -p` results) and how to run outside the checkout. The lexical floor is not extended and no model-graded follow-up is filed without the owner's answer. - **CHANGELOG (#3612, reverse).** Deleted the `### Recorded` interop bullets and the fleet-scan sentence from the 0.24.7 entry: the operator parked those four questions and the agent answered them anyway, so they are not decisions. Removed the duplicated Check 1 `model` bullet from 0.24.10. Both in-place edits are named in the 0.24.13 entry. - **`setup` eval prompt** drops the em dash that the #4838 re-serialization escaped as a unicode sequence (cross-group request from work-items). - **Release.** Manifest 0.24.12 to 0.24.13 with one CHANGELOG entry. ## Verification - `bash plugins/skill-quality/scripts/check-skill.test.sh`, `check-evals-quality.test.sh`, `check-listing-budget.test.sh`, `measure-invocation.test.sh`: all pass. - `bash plugins/skill-quality/scripts/check-skill.sh` over the `check` skill: PASS, 0 errors, 0 warnings. - `bash scripts/check-changelog-parity.sh --check --check-order`: clean. - `bash scripts/validate-plugins.sh`: all manifests and the catalog validated. - `bash scripts/check-docs-naming.sh`: clean. - `markdownlint-cli2` on the edited markdown (CHANGELOG, README, SKILL.md, invocation-probes.md): 0 issues. - Merged `origin/main` into the branch; no conflicts. ## Related - Audit findings: `.work/audit/REPORT.md` rows for #4586 (3c, evals.json:108), #3612 (3b reopen and 3d reverse), #3526 (3b accept, owner narrowed). - Issue operations, done outside this PR: reopen #3612 with the four interop questions as a decision packet and the findings report (`.work/fix/artifacts/skill-quality/3612-findings.md`); a decision packet on #3526 (is the lexical floor enough, CI drift, placement). - Cross-group requests: `scripts` group, add a `check-changelog-parity.sh` rule that fails a newly added entry whose body duplicates another entry in the same CHANGELOG (the 0.24.9 and 0.24.10 case recurs in autonomy, claude-config, claude-ops, context-budget and playbooks). `claude-config` group, mention this PR's check 1 `model` relaxation as a follow-up when commenting on #4027. - Cross-group requests from `tracker` and `decisions-docs` about #3612: already handled. #3612 is reopened with the four interop questions and a comment that checks #5045 (it kept only the spec-mandated check 1 changes; its `### Recorded` answers are removed here), so no further action. - Not touched: `plugins/skill-quality/hooks/exec-bash.mjs`; seed probes stay in the plugin because CI validates them. 🤖 Generated with [Claude Code](https://claude.com/claude-code) https://claude.ai/code/session_01EugXnFddtpHcY5gTuyEirB --------- Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
kyle-sexton
added a commit
that referenced
this pull request
Sep 30, 2026
…-limit description (#5512) Closes #3612 ## Summary The four Agent Skills interop questions on #3612 are decided, and the decisions had no durable home. This records all four as the owner's 2026-09-29 decisions on #3612 (option A on each), and fixes the last skill description over the 1024-codepoint limit that the description-cap baseline still tolerated. ## Fix - Adds an ADR recording Q1 to Q4 with reasons and the per-rule gate summary, and points the `.agents` rows in `docs/specs/agent-doc-surfaces.md` at it instead of the stale pending wording. - Trims the `work-items:setup` description under 1024 codepoints, keeping its trigger phrases, and drops its entry from `scripts/skill-description-cap-baseline.txt` (the baseline is now empty of over-limit skills). - `improvement:find` is not touched here; its trim is #5422. ## Verification - After merging origin/main: `scripts/check-changed-skills.sh origin/main` (setup PASS, 0 errors), `scripts/check-changelog-parity.sh --check --check-order` and `--check-bump origin/main`, and `scripts/validate-plugins.sh` all pass. work-items is 0.42.1, above main's 0.42.0. - `scripts/affected-tests.sh --run`: 45 pass, 2 fail. Both failures are in `check-script-contract.test.sh` and come from `node_modules/.bin/htmlhint` being absent in this worktree (`npm ci` not run); no touched file is involved. - `ci-status` on this PR. ## Related #5359, #5307, #5045, #5422 (improvement:find trim), #2911 🤖 Generated with [Claude Code](https://claude.com/claude-code) --------- Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Refs #3612.
Summary
Check 1 fails a plain
descriptioncontaining": "(YAML mapping indicator — unparsed frontmatter loads the skill with no fields set) and fails a presentcompatibilityoutside 1–500 characters. Quoted and block scalars stay valid. Absence ofcompatibilitystays valid.plugins/skill-quality→0.24.7.Does not use Closes — the issue was closed by a park ledger merge; reopen is blocked for this token (403). Pair with the ledger revert PR.
Tests
check-skill.test.sh: all assertions passed.