test(evals): three-case floor for fifteen thin skills (#4070) - #4838
Merged
Merged
Conversation
Each listed skill now ships three or more realistic evals in the runner shape. The skill-quality case-count advisory stays advisory. Closes #4070 Co-authored-by: Kyle Sexton <kyle-sexton@users.noreply.github.com>
…n-flow to 0.38.7 Co-authored-by: Kyle Sexton <kyle-sexton@users.noreply.github.com>
…playbooks session-flow 0.38.8, testing 0.9.4, playbooks 0.13.10 sit above in-flight 0.38.7 / 0.9.3 / 0.13.9 so changelog-parity does not collide on merge. Co-authored-by: Kyle Sexton <kyle-sexton@users.noreply.github.com>
Contributor
|
PR body contract — issue linkage This PR body does not yet satisfy the issue-linkage contract:
Edit the body and this comment updates itself on the next run. |
…l-floor-37e9 Co-authored-by: Kyle Sexton <kyle-sexton@users.noreply.github.com>
…l-floor-37e9 Co-authored-by: Kyle Sexton <kyle-sexton@users.noreply.github.com>
Co-authored-by: Kyle Sexton <kyle-sexton@users.noreply.github.com>
This was referenced Sep 29, 2026
kyle-sexton
added a commit
that referenced
this pull request
Sep 29, 2026
…s from boris evals The hygiene catalog's residue-dissolve note pointed at dissolve-comments' exclusion twice; it now states the example once and names the owning section. Boris eval cases 1 and 3 lose the escaped em dashes left by the #4838 re-serialization. The unreviewed boris amendment from #5154 was checked against the model-config page and stands. Extends the 0.14.0 entry. Refs #4027 Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01EugXnFddtpHcY5gTuyEirB
This was referenced Sep 29, 2026
Merged
kyle-sexton
added a commit
that referenced
this pull request
Sep 29, 2026
No related issue: cosmetic follow-up from the #4838 audit (unit 4070), routed from the work-items fix group. ## Summary `plugins/miro/skills/setup/evals/evals.json` carried a `—` escape in the case-3 prompt, introduced when #4838 re-serialized the file. ## Fix Remove the em dash from the prompt: `/miro:setup Miro tools are missing. Is my token wrong?`. Bump `miro` to 0.4.16 with a changelog entry. ## Verification - `jq` parses the file; the expectations and case text are unchanged. - `scripts/check-changelog-parity.sh --check`, `--check-bump origin/main`, `--check-order`: pass. - `scripts/validate-plugins.sh`: pass. ## Related Audit 2026-09-27..29, group miro. The other miro finding (4144, dependabot dist note) belongs to `scripts/dependabot-plugin-bump.sh` and is routed to the scripts group. 🤖 Generated with [Claude Code](https://claude.com/claude-code) https://claude.ai/code/session_01EugXnFddtpHcY5gTuyEirB Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
kyle-sexton
added a commit
that referenced
this pull request
Sep 29, 2026
…dit follow-ups (#5295) Closes #4526 Closes #4575 Refs #4001, #4027, #4346, #4532, #4533, #4535, #4574, #4601, #4537, #4503, #4504, #5167 ## Summary Fixes the playbooks-group findings from the audit of the 2026-09-27 to 2026-09-29 Cursor agent PRs. The changes are in `repo-sweep` (partial coverage, a declined outcome, review and prepare order, the hygiene catalog, version-stamped records), `skill-authoring` (argument guidance), and `boris` (effort amendment). #4526 and #4575 are reopened by the audit and met in full on this branch, so this PR closes them. #4346 stays open: it needs a Fable 5.1 session and cannot be done on another model. ## Fix - `history.sh`: partial coverage is decided from the newest merged sweep PR that names the step, so one old partial tick no longer forces a rerun forever; a `, not applicable:` line is not a mention (#4526, #4533). - `tick.sh`, `state.sh`, `SKILL.md`: `--partial <detail>` works with `committed` and `report-only`; `SKILL.md` Formats is the single list of done-line forms (#4526). - `tick.sh`, `state.sh`, `next.md`: `declined <n>` outcome; scope decisions are posted on the sweep PR when no commit is made (#4575). - `review.md`, `next.md`: ask the user before dispatching the reviewer; four routing rows for the four classes; re-check applies-when before priming; prime column named as column 8 (#4574, #4532, #4533). - `catalogs/hygiene.md`, `catalog.test.sh`: primes per entry, drops the #4503 and #4504 overrides, tests `prime: false` (#4532, #4535). - `SKILL.md` records: worktree-guard record and bundled-skill gotcha carry the Claude Code version they were rechecked against (#4601, #4537). - `skill-authoring`: argument guidance points at the skill argument shape convention (#4001). - `boris/reference/autonomy.md`: effort amendment re-verified against the model-config page (#4027). - `catalogs/hygiene.md`: the `residue-dissolve` note states the CI-workflow override example once and points at `dissolve-comments` Hard rules for the exclusion (cross-group request from ci). - `boris/evals/evals.json`: cases 1 and 3 drop the escaped em dashes left by the #4838 re-serialization (request from work-items, #4070). - `boris/reference/autonomy.md`: the unreviewed #5154 amendment was checked line by line against the model-config page (2026-09-29, 109,282 bytes) and stands; only a long line was reflowed (request from claude-config, #4027). - Version 0.14.0 and one CHANGELOG entry. Released-entry edits, named per the changelog-parity rule (#2388), and also named in the 0.14.0 entry: - `plugins/playbooks/CHANGELOG.md` 0.13.22: the body repeated 0.13.19 and 0.13.21; replaced with a statement that the release changed no plugin content. - `plugins/playbooks/CHANGELOG.md` 0.13.12: added a sentence that the named-slots pilot was reversed in 0.13.29. ## Verification - `scripts/check-changelog-parity.sh --check --check-order`: pass. - `scripts/check-changelog-parity.sh --check-preserved origin/main`: pass (101 headings). - `scripts/validate-plugins.sh`: all manifests and the catalog validated. - All nine `*.test.sh` under `plugins/playbooks` (repo-sweep: history, tick, state, catalog, render, guard, skill-version; boris update; skill-authoring update): pass. - Not done: a live worktree-isolated probe of the three refused Bash forms; the SKILL.md record says "not re-probed". ## Related Findings are in `.work/audit/REPORT.md` (local audit), by issue: #4526 (rows 3b and 3c, history.sh), #4575 (3b), #4346 (3b, partial delivery; stays open, needs-human, Fable 5.1 session), #4532, #4533, #4535, #4574 (3c review.md, 3c catalog.test.sh), #4601, #4537, #4001, #4027. Cross-group requests: - conventions: restore the owner-adopted wording of ground 2 in `docs/conventions/skill-argument-shape/README.md` ("It is destructive, such as --execute or --force"), per #4001. - claude-config: update `docs/upstream/claude-code.md` line 30 and `plugins/claude-config/skills/audit-instructions/reference/criteria.md` I21 to the 2026-09-28 model-config read (follows the boris re-verification). - ci: make a PR body containing "Do not merge" fail `ci-status` or apply `do-not-merge` (process failure behind playbooks 0.13.17, #5154). Cross-group requests received and handled here: - work-items, ci, claude-config: applied (see Fix). - docs-hygiene (duplicate 0.13.19/0.13.22 CHANGELOG text): already done by the 0.13.22 trim above; nothing more to change. - code-tidying (hygiene catalog overrides, in-place, staged wording): already done in `hygiene.md`; `catalog.test.sh` keeps its `issue` fixture because `issue` is a documented optional catalog key the parser still reads. - planning (`next.md` line 85 and the `/planning:interview scope` action): not edited; it waits on the owner answer to the #4502 decision packet. 🤖 Generated with [Claude Code](https://claude.com/claude-code) https://claude.ai/code/session_01EugXnFddtpHcY5gTuyEirB --------- Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
kyle-sexton
added a commit
that referenced
this pull request
Sep 29, 2026
…nt (#5243) No related issue: audit finding fix ## Summary Audit finding `plugin-education` (changelog-integrity, low): the `education` 0.11.6 changelog entry claimed more than the change did, the 0.11.3 entry ended in a bare `(#4119)` instead of a link, and the `teach` argument hint had dropped its action list. Ships as `education` 0.11.7. ## Fix - `plugins/education/CHANGELOG.md`: `[0.11.6]` now says only that the examples moved into the skill body (the old hint carried only `(e.g., ...)` examples); `[0.11.3]` links `#4119` like 0.11.4 to 0.11.6; new `[0.11.7]` entry. - `plugins/education/skills/teach/SKILL.md`: `argument-hint` lists the closed action set from the Action Router table (95 characters, inside the 100-character budget). - `plugins/education/.claude-plugin/plugin.json`: version 0.11.6 to 0.11.7. - `plugins/education/skills/quiz-me/evals/evals.json` case 9 lists a new auth-middleware diff fixture (`evals/fixtures/auth-middleware-change/auth-middleware.diff`: public-path guard, rate-limit-before-auth ordering, invalid-token 401 path) and names those behaviors in its first expectation, so the sandbox has a change to quiz on (planning request, #3589 AC2). - `plugins/education/skills/setup/evals/evals.json` case 3: the escaped em dash in the prompt is replaced by a colon (work-items request, F10 of unit 4070). - In-place corrections to released entries: `[0.11.6]` reworded to say only that the examples moved into the skill body; `[0.11.3]` issue reference linked. Both are repeated in the `[0.11.7]` entry, as `scripts/check-changelog-parity.sh` requires. ## Verification - `bash scripts/check-changelog-parity.sh --check --check-order`: pass. - `bash scripts/check-changelog-parity.sh --check-bump origin/main`: pass. - `bash scripts/check-changelog-parity.sh --check-preserved origin/main`: pass (45 headings compared). - `bash scripts/validate-plugins.sh`: all manifests and the catalog validated. - `bash scripts/validate-plugin-contracts.test.sh`: PASS=92 FAIL=0, zero argument-hint warnings in the shipping tree. - `education` has no test scripts of its own. ## Related - Audit report: `.work/audit/REPORT.md`, finding `plugin-education`. - #3542 (argument-hint house style, delivered earlier), #4119 (linked in the 0.11.3 entry). - Cross-group requests applied: work-items F10 (unit 4070, #4838) and planning quiz-me case 9 (#3589 AC2). A paid with/without eval run of case 9 stays with the owner, so #3589 is Refs only. - Verification of both: the two evals.json files parse, `plugins/skill-quality/scripts/check-evals-quality.sh` passes on each, and the changelog parity checks pass. 🤖 Generated with [Claude Code](https://claude.com/claude-code) https://claude.ai/code/session_01EugXnFddtpHcY5gTuyEirB --------- Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
kyle-sexton
added a commit
that referenced
this pull request
Sep 29, 2026
…-invocation portability (#5307) Refs: #3526 Refs: #3612 Refs: #4586 ## Summary Fixes the skill-quality findings from the audit of the unattended Cursor agent's PRs (REPORT.md rows for #4586, #3612, #3526). `check` documentation and its own eval case now agree that check 25 is a blocking FAIL. Check 1 no longer hard-errors on provider-format `model` ids and now reads the whole `description` scalar. `measure-invocation score` fails clearly outside the marketplace checkout. The agent's unratified interop bullets are removed from the CHANGELOG. Ships as `skill-quality` 0.24.13. No `Closes`: #4586 and #3526 are already closed records, and #3612 keeps open acceptance (the operator's four interop answers). ## Fix - **Check 25 is a FAIL (#4586).** `evals.json` case 9 graded an advisory WARN, the opposite of `check-skill.sh`. README, the `check` step 3 output guide and the `check` description no longer call the polarity check advisory or say the skill never blocks. The manifest description names the `measure-invocation` harness. - **`check` gets `## Next`** and drops a paragraph that repeated the Arguments text (check 27b self-trip). - **Check 1, `model` (#4027, #3612).** Fails only an empty or whitespace-containing value; Bedrock ids with `:`, ARNs and Vertex ids with `@` pass. The skills page defines no stricter grammar. - **Check 1, `description` (#3612).** The unquoted-colon check covers every continuation line of a plain scalar and a line ending in `:`. Quoted and block scalars stay exempt. - **`measure-invocation` (#3526).** `score` exits 2 with guidance when no probe `skill_dir` resolves; `validate` adds an unresolved-count line; the `check` action resolves the probes directory explicitly. `reference/invocation-probes.md` states the lexical floor's limits (seed positives quote the description, no ingest path for plugin-eval or `claude -p` results) and how to run outside the checkout. The lexical floor is not extended and no model-graded follow-up is filed without the owner's answer. - **CHANGELOG (#3612, reverse).** Deleted the `### Recorded` interop bullets and the fleet-scan sentence from the 0.24.7 entry: the operator parked those four questions and the agent answered them anyway, so they are not decisions. Removed the duplicated Check 1 `model` bullet from 0.24.10. Both in-place edits are named in the 0.24.13 entry. - **`setup` eval prompt** drops the em dash that the #4838 re-serialization escaped as a unicode sequence (cross-group request from work-items). - **Release.** Manifest 0.24.12 to 0.24.13 with one CHANGELOG entry. ## Verification - `bash plugins/skill-quality/scripts/check-skill.test.sh`, `check-evals-quality.test.sh`, `check-listing-budget.test.sh`, `measure-invocation.test.sh`: all pass. - `bash plugins/skill-quality/scripts/check-skill.sh` over the `check` skill: PASS, 0 errors, 0 warnings. - `bash scripts/check-changelog-parity.sh --check --check-order`: clean. - `bash scripts/validate-plugins.sh`: all manifests and the catalog validated. - `bash scripts/check-docs-naming.sh`: clean. - `markdownlint-cli2` on the edited markdown (CHANGELOG, README, SKILL.md, invocation-probes.md): 0 issues. - Merged `origin/main` into the branch; no conflicts. ## Related - Audit findings: `.work/audit/REPORT.md` rows for #4586 (3c, evals.json:108), #3612 (3b reopen and 3d reverse), #3526 (3b accept, owner narrowed). - Issue operations, done outside this PR: reopen #3612 with the four interop questions as a decision packet and the findings report (`.work/fix/artifacts/skill-quality/3612-findings.md`); a decision packet on #3526 (is the lexical floor enough, CI drift, placement). - Cross-group requests: `scripts` group, add a `check-changelog-parity.sh` rule that fails a newly added entry whose body duplicates another entry in the same CHANGELOG (the 0.24.9 and 0.24.10 case recurs in autonomy, claude-config, claude-ops, context-budget and playbooks). `claude-config` group, mention this PR's check 1 `model` relaxation as a follow-up when commenting on #4027. - Cross-group requests from `tracker` and `decisions-docs` about #3612: already handled. #3612 is reopened with the four interop questions and a comment that checks #5045 (it kept only the spec-mandated check 1 changes; its `### Recorded` answers are removed here), so no further action. - Not touched: `plugins/skill-quality/hooks/exec-bash.mjs`; seed probes stay in the plugin because CI validates them. 🤖 Generated with [Claude Code](https://claude.com/claude-code) https://claude.ai/code/session_01EugXnFddtpHcY5gTuyEirB --------- Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
kyle-sexton
added a commit
that referenced
this pull request
Sep 29, 2026
…ce (#5265) Refs: #4052 Refs: #3542 ## Summary Audit fixes for the `testing` plugin. #4052 was closed by #4828 with the `run-e2e` Native step only partly written; this delivers the missing wrap grammar, removes the contradiction with the Boundary section, and states Arguments once in the four skills that repeated them. ## Fix - `run-e2e`: the Native step for the bundled `run` skill now has the full wrap grammar from `docs/conventions/native-references/README.md`: gate token, identity check (a project skill named `run` is a legitimate target, not a skip), mutation fingerprint, skip and refuse states naming the axis line and enable path, a result block, and a rule that an unattended run never invokes `run`. - `run-e2e`: one launch path per verification. An orchestrator-governed start skips the Native step; the Boundary section and `context/bundled-run.md` now match. This is a judgment call for owner review. - `diagnose`, `plan`, `write`, `run-e2e`: the top `**Arguments.**` line is folded into the `## Arguments` section, with examples. - `run-e2e/evals/evals.json`: the case that duplicated case 1 is replaced; four cases cover the Native step. - `diagnose`, `plan`, `write` evals: `\u2014` escapes from the #4838 re-serialization are literal em dashes again (parsed JSON identical; `run-e2e` was already clean). Cross-group request from work-items (F10, #4070). - `testing` 0.9.7 to 0.10.0, CHANGELOG entry. - Not changed: the description opening with routing text, which the convention requires (README lines 56-60). The store row (`integration=wrap`, `baked.native_step=true`) is untouched. ## Verification - `scripts/check-changelog-parity.sh --check --check-order`: passed. - `scripts/validate-plugins.sh`: all manifests and the catalog validated. - `scripts/check-changed-skills.sh origin/main`: 4 skills checked (diagnose, plan, run-e2e, write), 0 failed, 0 warnings. - `bash plugins/testing/skills/audit/scripts/cant-fail-scan.test.sh`: 296 checks passed. - Not run: #4052 AC3 (live positive and negative transcripts) needs a host where the bundled `run` resolves and one with `disableBundledSkills`. `overlap.py self-check` exits 3 (degraded: recorded extraction versions differ from build 2.1.284), not a literal 0. ## Related - `.work/audit/REPORT.md` row 3b #4052 (reopen), finding on #3542 (Arguments stated twice, testing share), #4070 (eval content only, issue closed). - #4052 is reopened with the exact operator steps for AC3 and the exit-3 note for AC2; the PR uses `Refs`, not `Closes`, because AC2 and AC3 are not literally met. - Cross-group, applied: work-items F10 (#4070, #4838 escapes) in the three evals above. Not assigned: `plugins/dometrain` carries the same escapes (grounding, setup, sync) and has no fix group. - Cross-group: claude-ops owns #3542 and the overlap store re-derive (comment the new exit code on #4052 if re-derived). conventions: `docs/conventions/native-references/README.md` line 218 Adopters row for `/testing:run-e2e` should now list description phrase, Boundary and Native step (outside `plugins/testing`). 🤖 Generated with [Claude Code](https://claude.com/claude-code) https://claude.ai/code/session_01EugXnFddtpHcY5gTuyEirB --------- Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
kyle-sexton
added a commit
that referenced
this pull request
Sep 29, 2026
…e verb, and skill docs (#5325) No related issue: audit remediation; no owned issue has every acceptance criterion met on this branch ## Summary Fixes the work-items plugin audit findings. Refs only, no `Closes`: #4598 awaits an owner decision, and the other issues have follow-ups outside this plugin or an open bound question. - #4598: the work-loop background-job paragraph stated a harness limitation as fact and had no verification record. It now reports the observation as interim, conditional on session isolation, with a four-part record. No option is chosen; the decision stays with the owner. - #4609 / #4690: the onboard-adapter generator refused specs carrying the seam's `release` verb. It now generates it. - #4610: the work-loop in-flight exclusion reported only a PR number. It now reports draft state and age, and the skill states that the exclusion has no age bound. - #4262: the concurrency-cap wording said an unset cap meant an internal default; `implement-dispatch` owns the cap and it does not bind under worker authority. - #4605, #4657: triage exits point at the lane-barred branch; the `0.41.1` entry states which trigger phrases were dropped. - #4051: argument hints are grammar-only; one arguments statement per skill. - Eight skills gain a `## Next` section. ## Fix - `skills/work-loop/SKILL.md`, `reference/escalation-marker.md`: interim paragraph plus verification record; escalation-marker sections back in writer/reader order. - `skills/onboard-adapter/scripts/generate-adapter.sh` and its test, template, spec, gitea and linear `adapter-spec.json`, `tracker-seam.md`, `providers.md`: `release` added to the verb set; `features.leases` required only when `release` is true; drift test pins the verb set to the dispatcher's public verbs plus `list-items`. - `tools/work-item-tracker/adapters/github/README.md` and `skills/work-loop/reference/mode-drain.md`: reporting reduction emitting `{number, isDraft, createdAt}`; report line `in flight: #<item> (PR #<pr>, draft|ready, open <age>)`. - `.claude-plugin/plugin.json`, `README.md`, `skills/work`, `skills/work-loop`: cap wording. - `skills/triage`, `CHANGELOG.md` (0.41.1 entry), `skills/scan-todos/evals/evals.json`, `skills/work/evals/evals.json`: triage pointer, changelog correction, escaped em dashes. - Eight `SKILL.md` files: `## Next`. - Cross-group requests applied: the wave-cap wording in `README.md`, `skills/work/SKILL.md`, `.claude-plugin/plugin.json`, and `skills/work-loop/SKILL.md` names the `implement_dispatch_wave_cap` operator option in the precedence and says the cap changes behavior only under commit authority `orchestrator` (options block regenerated with `scripts/sync-plugin-options-docs.py`); `skills/scan-todos/evals/evals.json` case 3 prompts `--work` as the skill defines it (auto-select the smallest group). - Cross-group requests not applied as code: the tracker REST path for `create-item` and the `add` pre-flight is filed as #5338 (Refs #3169); the babysit-loop pointer landed in #5317, and #4598 got a comment linking it. - Version 0.41.8 to 0.41.9 (patch) with one changelog entry. `plugins/work-items/hooks/exec-bash.mjs` is untouched. ## Verification - `git merge origin/main`: clean, no conflicts. - `scripts/check-changelog-parity.sh --check --check-order`: passes. - `scripts/validate-plugins.sh`: all manifests and the catalog validate. - Every `*.test.sh` under `plugins/work-items` (about 70 suites, including `generate-adapter.test.sh`, the adapter capabilities tests, and the conformance suites): all exit 0. ## Related Refs #4598 Refs #4605 Refs #4609 Refs #4610 Refs #4070 Refs #4262 Refs #3169 (via #5338) Audit findings by issue: #4598 (F2, F3, F4, F8, F12), #4605 (F13, F14), #4609/#4690 (F5, F16), #4610 (F6, F17), #4262 (F7, F11), #4657 (F15, F18), #4051 (F20), #4070 (F10, F22), F23 (`## Next`). F14 and F21 need no edit (all six #4605 stranded items are closed; the skipped 0.41.3 was never released). Issue operations (reopen #4598 with a decision packet, comments on #4605, #4609, #4610, and an in-flight-bound follow-up under #4610) are separate from this diff. Cross-group requests: - core-docs: `docs/cloud-sessions.md` (create-item fails on GraphQL 403 in cloud sessions, links to the removed park ledger, "lease trio" wording). - decisions-docs: remove `docs/out-of-scope/cloud-session-permission-floor.md`. - conventions: `docs/conventions/loop-lane/README.md` background-job launch mode paragraph, to point at the work-loop skill. - implementation: retire-or-keep for `work_dispatch_concurrency_cap` (#4262, #4272 packet). - songwriting: `scan-todos` eval case 3 prompts `--work` as a filing request (#4070); applied. - tracker: REST path for `create-item` and the `add` pre-flight on gh 2.45 GraphQL-403 sessions; filed as #5338 (Refs #3169), not delivered here. - source-control: babysit-loop pointer to the loop-lane background-job paragraph merged in #5317; comment on #4598 links it. - implementation: wave-cap precedence and composed-budget wording; applied (names `implement_dispatch_wave_cap`). - education, miro, playbooks, session-flow, skill-quality, songwriting, testing: cosmetic `—` escapes in evals.json from the #4838 re-serialization (#4070). 🤖 Generated with [Claude Code](https://claude.com/claude-code) https://claude.ai/code/session_01EugXnFddtpHcY5gTuyEirB --------- Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Closes #4070
Summary
Fifteen skills that shipped one or two eval cases now each ship three or more in the runner shape. The skill-quality case-count advisory stays advisory (ADR 0023); this is coverage, not a new gate.
Fix
Added realistic cases (trigger/routing, a second happy path or empty-input, a refusal/guardrail) to:
dometraingrounding, setup, synceducationsetup,mirosetupplaybooksborissession-flowsetup,skill-qualitysetupsongwritingco-write, rhymetestingdiagnose, plan, run-e2e, writework-itemsscan-todosDecision
check-evals-quality.shexplicitly does not check case count (recorded divergence from volume-over-polish until the deferred runner lands).Verification
evals.jsonhas ≥3 cases.check-evals-quality.sh: PASS (4 pre-existing Q4 WARNs).session-flow0.38.7: main already claimed 0.38.6).