-
Notifications
You must be signed in to change notification settings - Fork 2
skill-quality: measure and improve description-driven auto-invocation (eval-based) #3526
Copy link
Copy link
Closed
Copy link
Labels
agent-readyFully specified and briefed; eligible for autonomous pickup from the frontier.Fully specified and briefed; eligible for autonomous pickup from the frontier.priority: mediumReal value, no hard deadline; normal backlog flow.Real value, no hard deadline; normal backlog flow.status: readyTriaged, unblocked, and fully specified; eligible to pick up.Triaged, unblocked, and fully specified; eligible to pick up.work-class: scopedA briefed fix or small feature; blast radius bounded by the brief, tests exist.A briefed fix or small feature; blast radius bounded by the brief, tests exist.
Description
Activity
Metadata
Metadata
Assignees
Labels
agent-readyFully specified and briefed; eligible for autonomous pickup from the frontier.Fully specified and briefed; eligible for autonomous pickup from the frontier.priority: mediumReal value, no hard deadline; normal backlog flow.Real value, no hard deadline; normal backlog flow.status: readyTriaged, unblocked, and fully specified; eligible to pick up.Triaged, unblocked, and fully specified; eligible to pick up.work-class: scopedA briefed fix or small feature; blast radius bounded by the brief, tests exist.A briefed fix or small feature; blast radius bounded by the brief, tests exist.
Motivation
The frontmatter-alignment interview (#3524) settled that discoverability rides almost entirely on
descriptioncontent: auto-invocation is the model matching a request against the listing text, so description quality is the fleet's highest-leverage unmeasured surface. The fleet currently passes all static checks (zero skills over the 1,536 cap, trigger phrases enforced bycheck-skill.sh), but nothing measures whether the descriptions actually win the invocations they should.One question is already closed, by citation
Packed description vs
when_to_usesplit: nothing to measure. The official frontmatter reference stateswhen_to_useis "Appended todescriptionin the skill listing and counts toward the 1,536-character cap", and the harness joins them with a literal 3-char " - " (encoded incheck-skill.sh's combined-length check andaudit_skill_visibility.py'sListingConfig.joiner_chars). Identical content split differently produces byte-near-identical model input, so an A/B on the field split has no effect to detect. The fleet keeps the packed pattern; this issue is about the content, not the field layout.Proposed work
claude plugin evalfixtures and/or headlessclaude -pruns, where each probe is a realistic user request and the grade is whether the intended skill is selected (and no competitor is). Start with a small probe set over the skills with the weakest invocation evidence (claude-ops:audit-skill-visibility's starvation report is the ranking input).check-skill.sh's trigger-preservation constraint; re-grade.pathsadoption for genuinely file-scoped skills (a precision lever; wrong globs suppress, so adopt only with probe coverage).claude-ops:audit-performanceat 1,526 andclaude-ops:audit-skill-visibilityat 1,518 of 1,536): trim only if probes show tail content matters.when_to_useconvergence for the 3 outlier skills using it (cosmetic; only if free).Constraints
check-skill.shfails dropped triggers vs base).claude-ops:audit-skill-visibility; this issue measures per-skill description quality only.