diff --git a/docs/CATALOG.md b/docs/CATALOG.md index 31749cd8be..b681769a54 100644 --- a/docs/CATALOG.md +++ b/docs/CATALOG.md @@ -81,7 +81,7 @@ plugin manifests and kept in sync by CI — never hand-edit it; the category voc - [`playbooks`](../plugins/playbooks) — Doctrine and knowledge playbooks as on-demand skills, plus a maintainer-facing update skill. boris — Boris Cherny's Claude Code workflow tips (howborisusesclaudecode.com); skill-authoring — Anthropic's internal skill-authoring playbook; fable-5 — Claude Fable 5's operating doctrine (self-authored, no upstream). The boris and skill-authoring packs vendor a verbatim upstream baseline; /playbooks:update drift-checks and syncs those baselines centrally (maintainers). - [`claude-config`](../plugins/claude-config) — Nine configuration-health skills (plus setup) for a repo's Claude Code configuration: audit (settings.json / .mcp.json / hooks / plugins / permissions drift), audit-automation-gaps (evidence-gated verdicts on automation gaps), audit-permission-grants (allow-rule / allowed-tools grants for auto-mode durability and portability), audit-permission-state (the permission rules actually in effect — every settings scope merged with per-rule provenance, what auto mode drops on entry, config written where nothing reads it, and which managed intents are enforced versus loosenable), draft-auto-mode-rules (interview and draft a paste-ready autoMode classifier block; prints only, never writes), audit-instructions (locally-owned instruction surfaces vs current model capability — proposes removals/rewrites of instructions the model no longer needs, and detects cross-surface instruction conflicts), audit-prompting-postures (the additive lane — posture guidance the prompting guide says a component's purpose needs but the component does not carry), audit-pass (one coordinated, ordered, resumable pass over a named target — three-scope inventory, run-time-derived exclusion set, stable finding identity, suppression memory, resume, one human gate — delegating every check to the plugin that owns it), and unhobble (the empirical bare-baseline experiment: reversibly strip a repo's standing instructions, log real stumbles against the current model, re-add only what evidence earns). - [`claude-memory`](../plugins/claude-memory) — Keeps a repo's Claude Code memory layer healthy and under your control, against criteria derived from official Claude Code documentation. The audit skill checks the instruction/memory layer (CLAUDE.md, CLAUDE.local.md, .claude/rules/, auto-memory) with a deterministic script-backed spine plus judgment-tier checks. The stateless skill inspects, disables, and (confirm-gated) purges Claude-written auto memory across all settings scopes. -- [`claude-ops`](../plugins/claude-ops) — Claude Code operations toolkit. Twelve skills: audit-skill-visibility (audit whether each installed skill is actually VISIBLE to the model, and diagnose why most of a fleet never gets used — a skill is invisible when its description is dropped by Claude Code's skill-listing context budget, which drops descriptions least-invoked-first so an unused skill loses the keywords that would let it be matched, from skills genuinely not wanted, from skills the run cannot observe at all; computes whether the listing overflows from documented settings, and withholds every cold verdict the data cannot support rather than reporting absence of data as absence of use), inventory (read-only enumeration of the complete invocable surface — every built-in CLI command with aliases and hidden/gated status, every bundled skill, and every component of every installed plugin across all marketplaces; reads the shipped binary because upstream publishes no built-in command list, and carries an integrity verdict so a drifted build reports counts as floors rather than silently short totals), audit-install-state (read-only audit of the machine-scope ~/.claude installation directory and ~/.claude.json — full inventory split into an authored surface and rolled-up bulk trees, product-managed retention vs genuinely unmanaged state, filename-scheme resolution before any process-liveness check, and deliberate/mid-experiment detection; reports, never deletes), audit-performance (read-only slowness-diagnostic capture run at the moment the machine or a session feels slow: CLI version, retention-sweep health including the silent unparsable-settings pause, a timed census walk of the install tree as a sweep-cost proxy, active-session and plugin-fleet counts, a process census, and the fan-out layer, which covers a load-labelled no-op spawn baseline, every hook that will fire bucketed per-tool-call versus per-turn with its invocation shape, the configured statusline, subagent concurrency and spawn-depth ceilings against documented defaults, whether running sessions predate the settings file they are judged by, and orphan attribution by parent liveness rather than age; read against a bundled known-performance-issues reference that also records the causes tested and cleared; separates the four documented suspects of accumulated state, version regression, component bloat, and per-spawn fan-out cost, and routes remediation out; reports, never mutates, and never executes a discovered hook or statusline command), audit-native-overlap (map native Claude Code surfaces — built-in CLI commands, bundled skills, plugin-backed built-ins, session-provided skills — against the current repo's plugin skills and agents, so a custom component never silently duplicates what Claude Code itself ships; bare invocation is a read-only overlap report carrying the extraction's integrity floors and a shared-listing-budget exposure section, verdicts are human-gated in a committed store rendered into a generated registry whose every row carries an observable recheck trigger, and only an explicit apply step bakes presence-gated native references into descriptions and Boundary sections), observability (read locally captured telemetry — OTEL store, collector, hook-event JSONL, ccusage — with trend reports and store pruning), known-issues (search known Claude product GitHub bugs, check service health, maintain a persistent tracked-issue registry), changelog (ingest Claude Code changelog entries and integrate them into the current repo), plugins (bring a machine's plugin fleet current on demand — marketplace refresh, effective-scope updates including in-repo project/local installs, new-plugin install per policy, scope-divergence detection and explicit convergence), morning-brief (read-only gh-based operator morning view — queue-label counts, merge-ready PRs, parked decisions with their RECOMMENDED lines, and loop-lane telemetry freshness), lanes (start/restart/stop/status loop lanes as named background Claude Code sessions seeded from canonical prompt files, with per-lane model/effort, a repo-pull + marketplace-refresh launch step, and a consume-restarts action — an OS-schedulable reader that relaunches stopped lanes whose telemetry carries a restart_request), and a re-runnable setup action that settles where the known-issues registry lives. Plus a family of eight advisory *-audit hooks (API errors, config changes, instruction loads, permission denials, pre-compaction, skill usage, tool failures, and unsurfaced hook failures — the last also warns the user via systemMessage, since a hook that fails to launch enforces nothing and Claude Code surfaces the failure to nobody) that emit the shared hook-telemetry envelope, and a reference sink that maps envelopes into the hook-events.jsonl the observability skill reads. +- [`claude-ops`](../plugins/claude-ops) — Claude Code operations toolkit. Twelve skills: audit-skill-visibility (audit whether each installed skill is actually VISIBLE to the model, and diagnose why most of a fleet never gets used — a skill is invisible when its description is dropped by Claude Code's skill-listing context budget, which sheds descriptions lowest-score-first so an unused skill loses the keywords that would let it be matched, from skills genuinely not wanted, from skills the run cannot observe at all; computes whether the listing overflows from documented settings, and withholds every cold verdict the data cannot support rather than reporting absence of data as absence of use), inventory (read-only enumeration of the complete invocable surface — every built-in CLI command with aliases and hidden/gated status, every bundled skill, and every component of every installed plugin across all marketplaces; reads the shipped binary because upstream publishes no built-in command list, and carries an integrity verdict so a drifted build reports counts as floors rather than silently short totals), audit-install-state (read-only audit of the machine-scope ~/.claude installation directory and ~/.claude.json — full inventory split into an authored surface and rolled-up bulk trees, product-managed retention vs genuinely unmanaged state, filename-scheme resolution before any process-liveness check, and deliberate/mid-experiment detection; reports, never deletes), audit-performance (read-only slowness-diagnostic capture run at the moment the machine or a session feels slow: CLI version, retention-sweep health including the silent unparsable-settings pause, a timed census walk of the install tree as a sweep-cost proxy, active-session and plugin-fleet counts, a process census, and the fan-out layer, which covers a load-labelled no-op spawn baseline, every hook that will fire bucketed per-tool-call versus per-turn with its invocation shape, the configured statusline, subagent concurrency and spawn-depth ceilings against documented defaults, whether running sessions predate the settings file they are judged by, and orphan attribution by parent liveness rather than age; read against a bundled known-performance-issues reference that also records the causes tested and cleared; separates the four documented suspects of accumulated state, version regression, component bloat, and per-spawn fan-out cost, and routes remediation out; reports, never mutates, and never executes a discovered hook or statusline command), audit-native-overlap (map native Claude Code surfaces — built-in CLI commands, bundled skills, plugin-backed built-ins, session-provided skills — against the current repo's plugin skills and agents, so a custom component never silently duplicates what Claude Code itself ships; bare invocation is a read-only overlap report carrying the extraction's integrity floors and a shared-listing-budget exposure section, verdicts are human-gated in a committed store rendered into a generated registry whose every row carries an observable recheck trigger, and only an explicit apply step bakes presence-gated native references into descriptions and Boundary sections), observability (read locally captured telemetry — OTEL store, collector, hook-event JSONL, ccusage — with trend reports and store pruning), known-issues (search known Claude product GitHub bugs, check service health, maintain a persistent tracked-issue registry), changelog (ingest Claude Code changelog entries and integrate them into the current repo), plugins (bring a machine's plugin fleet current on demand — marketplace refresh, effective-scope updates including in-repo project/local installs, new-plugin install per policy, scope-divergence detection and explicit convergence), morning-brief (read-only gh-based operator morning view — queue-label counts, merge-ready PRs, parked decisions with their RECOMMENDED lines, and loop-lane telemetry freshness), lanes (start/restart/stop/status loop lanes as named background Claude Code sessions seeded from canonical prompt files, with per-lane model/effort, a repo-pull + marketplace-refresh launch step, and a consume-restarts action — an OS-schedulable reader that relaunches stopped lanes whose telemetry carries a restart_request), and a re-runnable setup action that settles where the known-issues registry lives. Plus a family of eight advisory *-audit hooks (API errors, config changes, instruction loads, permission denials, pre-compaction, skill usage, tool failures, and unsurfaced hook failures — the last also warns the user via systemMessage, since a hook that fails to launch enforces nothing and Claude Code surfaces the failure to nobody) that emit the shared hook-telemetry envelope, and a reference sink that maps envelopes into the hook-events.jsonl the observability skill reads. - [`rate-limit-guard`](../plugins/rate-limit-guard) — Shared rate-limit guard for loop lanes: a statusline wrapper tees the subscription rate-limit windows to a fixed machine-scope file, a StopFailure hook records rate-limit stops reactively, and a reader contract fixes how consuming sessions pause and resume. - [`context-guard`](../plugins/context-guard) — Per-session context-window observability plus the first shipped consumer: a statusline wrapper tees each session's context_window fields to a per-session snapshot file, a zone resolver classifies usage into smart/acceptable/dumb bands (percentage bands plus window-class token bands, conservative-min combination, zones.json SSOT with shipped defaults), a reader contract fixes how consuming sessions interpret the snapshots, and zone-crossing hooks report once per transition into a worse zone across two channels — the continuation menu to the operator, who owns that choice, and to the model only the zone determination plus the counter-steer that a zone word is not a decay signal (advisory by default; an optional blocking mode gates new mutating work on a fresh dumb-zone snapshot with handoff-writing exempt), with a PostCompact hook persisting an evidence-degraded marker. - [`context-budget`](../plugins/context-budget) — Measure a Claude Code session's fixed startup context payload per item, on the consumer's machine at a pinned, version-stamped binary — including per-tool attribution of the built-in tool pools that /context reports only as lump sums, derived live by A/B bare-name-deny differencing with enforced comparability rules (skill-listing signature, one mode, one binary), an SDK-primary exact meter degrading to a version-aware headless /context parser and then to an honest structured error (never a wrong number), and a per-project measure-toggle-remeasure ledger under the plugin data directory recording every lever's real before/after delta. Report-only: prints exact config, applies nothing. diff --git a/docs/NATIVE-SURFACES.md b/docs/NATIVE-SURFACES.md index 80c2cf3a61..349de0cfce 100644 --- a/docs/NATIVE-SURFACES.md +++ b/docs/NATIVE-SURFACES.md @@ -18,7 +18,7 @@ and when — see [`docs/conventions/native-references/`](conventions/native-refe | Lane | Rows | Baked | Verdicts | |---|---|---|---| | Built-in CLI commands | 1 | 0 | complementary 1 | -| Bundled skills | 6 | 1 | complementary 6 | +| Bundled skills | 7 | 1 | complementary 7 | | Plugin-backed built-ins | 1 | 0 | complementary 1 | | Session-provided skills (observation-only) | 1 | 0 | defer 1 | @@ -88,6 +88,21 @@ and when — see [`docs/conventions/native-references/`](conventions/native-refe - **Baked:** description phrase no · Boundary section no - **Budget caveat:** the baked phrase may be dropped from the skill listing under budget pressure — it is the best available routing surface, not a guaranteed one +### `doctor` → `claude-ops:audit-skill-visibility` + +- **Verdict:** `complementary` — Same native surface as the two sibling rows, a third of our lanes. Bundled `doctor` ships a one-shot check (its Check 1) that groups unused skills, MCP servers, and plugins against their context cost, labels each group with a token-savings estimate, and offers to disable the selected groups. audit-skill-visibility answers a different question, why a skill is unseen: it reconciles three usage sources (native ~/.claude.json counters, its own JSONL store, OTEL) under a max-across-sources rule, computes an observed horizon and withholds every verdict the span cannot support, diagnoses reachability causes, and analyses listing-budget starvation. It disables nothing by contract. The skill's own description and Scope boundary already route the one-shot unused-versus-context-cost question to the native surface; this row records that routing in the store rather than replacing it. +- **Native surface:** `doctor` (bundled skill; markers: gated) +- **Our component:** `claude-ops:audit-skill-visibility` (skill) +- **Evidence:** + - `doctor` present in the 2026-08-23 extraction as bundled-skill (markers: gated; aliases: checkup), per the two sibling rows + - the shipped doctor skill carries a check titled 'Check 1: unused skills, MCP servers, and plugins' whose prompt groups unused components, labels each group with a benefit estimate ('37 unused skills, saves ~2.2k est. tokens/session'), and applies only the groups the user selects; confirmed by string search of the installed v2.1.252 binary on 2026-08-31 + - our description: audit whether each installed skill is actually VISIBLE to the model; reconciles native counters, a JSONL store, and OTEL; withholds every verdict the data cannot support; read-only, never disables, deletes, or edits a skill + - our description's Not-for clause and the SKILL.md Scope boundary table both already name the native surface ('Claude Code ships that in /doctor and the Stats tab') with no store row behind them until this one; a prose disclaimer without a store row is the drift this registry exists to catch +- **Observation:** extraction — targeted string search of the installed binary v2.1.252 (doctor Check 1 strings confirmed; a spot observation over the sibling rows' full v2.1.232 extraction, not a re-extraction) (2026-08-31) +- **Recheck trigger:** a Claude Code release changes doctor's unused-components check (Check 1's grouping, its disable offer, or its benefit estimate), gives it a multi-source reconciliation or observation-horizon discipline, or changes /doctor's status as a bundled skill or its gating switch (verified 2026-08-31) +- **Baked:** description phrase no · Boundary section no +- **Budget caveat:** the baked phrase may be dropped from the skill listing under budget pressure — it is the best available routing surface, not a guaranteed one + ### `run` → `testing:run-e2e` - **Verdict:** `complementary` — The bundled skill answers 'did this change work when I ran the app'; run-e2e drives named UI and API flows, captures evidence (screenshots, responses, logs), and carries a non-UI smoke playbook for libraries, MCP servers, hooks, and scripts — surfaces that have no app to launch. Prefer the native surface for the quick look; ours where the verification has to be reproducible or the target is not an app. diff --git a/docs/adr/0016-source-skill-recommendation-from-the-catalog-not-the-listing.md b/docs/adr/0016-source-skill-recommendation-from-the-catalog-not-the-listing.md index fb8a5f206e..1cad962ce5 100644 --- a/docs/adr/0016-source-skill-recommendation-from-the-catalog-not-the-listing.md +++ b/docs/adr/0016-source-skill-recommendation-from-the-catalog-not-the-listing.md @@ -27,6 +27,30 @@ the wrong direction, and — the part that makes it a correctness bug rather tha cannot tell that it is blind. The gatekeeping the contract bans would have been reinstated by the harness, invisibly. +> **Revised 2026-08-31 ([#3534](https://github.com/melodic-software/claude-code-plugins/issues/3534)):** +> the drop-order sentence above restates ("Skill +> descriptions are cut short"), which is itself wrong and **still wrong as of 2026-08-31**. The error +> is upstream's, not this ADR's, which is why the same sentence keeps re-entering this repo: it is on +> `main` in the skill's own SKILL.md and in the plugin manifest, and #3524 added a citation pointing +> at the wrong page for it. Claude Code does not drop descriptions "starting with the skills invoked +> least". The shipped binary does something else, on two independent axes. +> +> It ranks by a decay-weighted score, `usageCount * max(0.5 ^ (daysSinceUse / 7), 0.1)`, so a +> heavily used but stale skill can be shed before a lightly used fresh one: 100 uses 21 days ago +> scores 12.5 and loses to 13 uses today. And it then walks every entry in that order with a running +> budget, granting whatever still fits, with no early exit, so a cheap never-invoked description can +> be granted after an expensive well-used one was refused. Description length is a second ranking +> input the documented account does not mention. Both were recovered from the 2.1.251 binary and +> re-verified at 2.1.252; the stamp, the greps, and the counterexamples are in +> [`plugins/claude-ops/skills/audit-skill-visibility/reference/listing-scorer.md`](../../plugins/claude-ops/skills/audit-skill-visibility/reference/listing-scorer.md). +> +> **The ADR's core decision is untouched, and this correction strengthens the case for it.** A +> never-invoked skill scores exactly zero and is still shed first, so the bias this paragraph +> identifies holds; the decay term adds a second bias the paragraph did not anticipate, against +> skills the operator used a while ago and has since forgotten, which is the same population +> `show-options` exists to surface. The budget arithmetic quoted above is unaffected: it measures +> demand against the budget, not the order of shedding. + Separately, the no-omission rule was measured against the real catalog at a real moment. Rendering every candidate in full produced **139 options across 275 lines, ~7 screens, ~4,100 tokens** — 97.8% of the catalog, i.e. the generated cheat sheet with an extra column, which an operator reads once and @@ -119,6 +143,33 @@ signals — and revisit the manual-only posture second. Usage-metrics-driven sur (`~/.claude.json` `skillUsage`, undocumented internal state) stays deferred; rotation runs off a ledger the skill writes itself, which is what keeps that deferral honest rather than load-bearing. +> **Revised 2026-08-31 ([#3534](https://github.com/melodic-software/claude-code-plugins/issues/3534)):** +> the deferral stands, but "undocumented internal state" is no longer the +> reason and should not be read as one. That substrate is now characterized and dated in +> [`plugins/claude-ops/skills/audit-skill-visibility/reference/usage-counters.md`](../../plugins/claude-ops/skills/audit-skill-visibility/reference/usage-counters.md), +> which invites the false inference that the deferral lifts once the state is known. It does not, +> because three grounds documentation cannot cure survive: +> +> - The scorer is **wrong-signed** for the question this skill asks. `zPe` scores high for skills +> used recently and often, which are exactly the skills the operator has not forgotten. Inverting +> it collapses to a decay-weighted least-recently-used ordering, which the self-written ledger +> already supplies without a build-pinned dependency. +> - `skillUsage` holds only skills that have fired at least once. It structurally cannot name the +> never-invoked population the no-omission rule above exists to protect. +> - It carries no causal-trigger field, so it cannot answer take-up: whether a skill was invoked +> *because* it was surfaced. The ledger can. That is not a stopgap for missing documentation; it +> measures something these counters never will. +> +> A binary-derived stamp is also not what "documented" meant here. Its own fallback rule, treat a +> version mismatch as scorer-unknown, is what ships alongside ground that can shift without notice, +> and its recheck trigger fires only when a human reads a release note. +> +> **Conditions for a future revisit,** so the next reader does not have to re-derive them: a +> published, versioned surface for skill-usage data rather than a reverse-engineered internal; a +> rotation signal not decay-weighted toward recent use; and take-up attribution. The third is +> unreachable from `skillUsage` by construction, so the ledger survives whatever happens to the +> first two. **The ADR's core decision is untouched.** + **The probe seam, and why the two-consumer version was withdrawn.** `show-options` adds no probe of its own — it routes to `orient` — and the duplication that decision sidestepped is resolved in the same change: `plugins/session-flow/reference/gather.md` now owns the block for all seven consumers diff --git a/docs/native-surfaces/records.json b/docs/native-surfaces/records.json index e77166f636..db48f20752 100644 --- a/docs/native-surfaces/records.json +++ b/docs/native-surfaces/records.json @@ -173,6 +173,33 @@ "baked": { "description_phrase": false, "boundary_section": false }, "budget_caveat": true }, + { + "native": { "name": "doctor", "class": "bundled-skill", "markers": ["gated"] }, + "component": { + "plugin": "claude-ops", + "skill": "audit-skill-visibility", + "kind": "skill" + }, + "verdict": "complementary", + "reason": "Same native surface as the two sibling rows, a third of our lanes. Bundled `doctor` ships a one-shot check (its Check 1) that groups unused skills, MCP servers, and plugins against their context cost, labels each group with a token-savings estimate, and offers to disable the selected groups. audit-skill-visibility answers a different question, why a skill is unseen: it reconciles three usage sources (native ~/.claude.json counters, its own JSONL store, OTEL) under a max-across-sources rule, computes an observed horizon and withholds every verdict the span cannot support, diagnoses reachability causes, and analyses listing-budget starvation. It disables nothing by contract. The skill's own description and Scope boundary already route the one-shot unused-versus-context-cost question to the native surface; this row records that routing in the store rather than replacing it.", + "evidence": [ + "`doctor` present in the 2026-08-23 extraction as bundled-skill (markers: gated; aliases: checkup), per the two sibling rows", + "the shipped doctor skill carries a check titled 'Check 1: unused skills, MCP servers, and plugins' whose prompt groups unused components, labels each group with a benefit estimate ('37 unused skills, saves ~2.2k est. tokens/session'), and applies only the groups the user selects; confirmed by string search of the installed v2.1.252 binary on 2026-08-31", + "our description: audit whether each installed skill is actually VISIBLE to the model; reconciles native counters, a JSONL store, and OTEL; withholds every verdict the data cannot support; read-only, never disables, deletes, or edits a skill", + "our description's Not-for clause and the SKILL.md Scope boundary table both already name the native surface ('Claude Code ships that in /doctor and the Stats tab') with no store row behind them until this one; a prose disclaimer without a store row is the drift this registry exists to catch" + ], + "observation": { + "class": "extraction", + "detail": "targeted string search of the installed binary v2.1.252 (doctor Check 1 strings confirmed; a spot observation over the sibling rows' full v2.1.232 extraction, not a re-extraction)", + "date": "2026-08-31" + }, + "recheck": { + "trigger": "a Claude Code release changes doctor's unused-components check (Check 1's grouping, its disable offer, or its benefit estimate), gives it a multi-source reconciliation or observation-horizon discipline, or changes /doctor's status as a bundled skill or its gating switch", + "verified": "2026-08-31" + }, + "baked": { "description_phrase": false, "boundary_section": false }, + "budget_caveat": true + }, { "native": { "name": "morning", "class": "session-skill", "markers": [] }, "component": { "plugin": "claude-ops", "skill": "morning-brief", "kind": "skill" }, diff --git a/plugins/claude-ops/.claude-plugin/plugin.json b/plugins/claude-ops/.claude-plugin/plugin.json index edf84fc30c..5d669e269d 100644 --- a/plugins/claude-ops/.claude-plugin/plugin.json +++ b/plugins/claude-ops/.claude-plugin/plugin.json @@ -1,8 +1,8 @@ { "$schema": "https://json.schemastore.org/claude-code-plugin-manifest.json", "name": "claude-ops", - "version": "0.38.23", - "description": "Claude Code operations toolkit. Twelve skills: audit-skill-visibility (audit whether each installed skill is actually VISIBLE to the model, and diagnose why most of a fleet never gets used \u2014 a skill is invisible when its description is dropped by Claude Code's skill-listing context budget, which drops descriptions least-invoked-first so an unused skill loses the keywords that would let it be matched, from skills genuinely not wanted, from skills the run cannot observe at all; computes whether the listing overflows from documented settings, and withholds every cold verdict the data cannot support rather than reporting absence of data as absence of use), inventory (read-only enumeration of the complete invocable surface \u2014 every built-in CLI command with aliases and hidden/gated status, every bundled skill, and every component of every installed plugin across all marketplaces; reads the shipped binary because upstream publishes no built-in command list, and carries an integrity verdict so a drifted build reports counts as floors rather than silently short totals), audit-install-state (read-only audit of the machine-scope ~/.claude installation directory and ~/.claude.json \u2014 full inventory split into an authored surface and rolled-up bulk trees, product-managed retention vs genuinely unmanaged state, filename-scheme resolution before any process-liveness check, and deliberate/mid-experiment detection; reports, never deletes), audit-performance (read-only slowness-diagnostic capture run at the moment the machine or a session feels slow: CLI version, retention-sweep health including the silent unparsable-settings pause, a timed census walk of the install tree as a sweep-cost proxy, active-session and plugin-fleet counts, a process census, and the fan-out layer, which covers a load-labelled no-op spawn baseline, every hook that will fire bucketed per-tool-call versus per-turn with its invocation shape, the configured statusline, subagent concurrency and spawn-depth ceilings against documented defaults, whether running sessions predate the settings file they are judged by, and orphan attribution by parent liveness rather than age; read against a bundled known-performance-issues reference that also records the causes tested and cleared; separates the four documented suspects of accumulated state, version regression, component bloat, and per-spawn fan-out cost, and routes remediation out; reports, never mutates, and never executes a discovered hook or statusline command), audit-native-overlap (map native Claude Code surfaces \u2014 built-in CLI commands, bundled skills, plugin-backed built-ins, session-provided skills \u2014 against the current repo's plugin skills and agents, so a custom component never silently duplicates what Claude Code itself ships; bare invocation is a read-only overlap report carrying the extraction's integrity floors and a shared-listing-budget exposure section, verdicts are human-gated in a committed store rendered into a generated registry whose every row carries an observable recheck trigger, and only an explicit apply step bakes presence-gated native references into descriptions and Boundary sections), observability (read locally captured telemetry \u2014 OTEL store, collector, hook-event JSONL, ccusage \u2014 with trend reports and store pruning), known-issues (search known Claude product GitHub bugs, check service health, maintain a persistent tracked-issue registry), changelog (ingest Claude Code changelog entries and integrate them into the current repo), plugins (bring a machine's plugin fleet current on demand \u2014 marketplace refresh, effective-scope updates including in-repo project/local installs, new-plugin install per policy, scope-divergence detection and explicit convergence), morning-brief (read-only gh-based operator morning view \u2014 queue-label counts, merge-ready PRs, parked decisions with their RECOMMENDED lines, and loop-lane telemetry freshness), lanes (start/restart/stop/status loop lanes as named background Claude Code sessions seeded from canonical prompt files, with per-lane model/effort, a repo-pull + marketplace-refresh launch step, and a consume-restarts action \u2014 an OS-schedulable reader that relaunches stopped lanes whose telemetry carries a restart_request), and a re-runnable setup action that settles where the known-issues registry lives. Plus a family of eight advisory *-audit hooks (API errors, config changes, instruction loads, permission denials, pre-compaction, skill usage, tool failures, and unsurfaced hook failures \u2014 the last also warns the user via systemMessage, since a hook that fails to launch enforces nothing and Claude Code surfaces the failure to nobody) that emit the shared hook-telemetry envelope, and a reference sink that maps envelopes into the hook-events.jsonl the observability skill reads.", + "version": "0.39.0", + "description": "Claude Code operations toolkit. Twelve skills: audit-skill-visibility (audit whether each installed skill is actually VISIBLE to the model, and diagnose why most of a fleet never gets used \u2014 a skill is invisible when its description is dropped by Claude Code's skill-listing context budget, which sheds descriptions lowest-score-first so an unused skill loses the keywords that would let it be matched, from skills genuinely not wanted, from skills the run cannot observe at all; computes whether the listing overflows from documented settings, and withholds every cold verdict the data cannot support rather than reporting absence of data as absence of use), inventory (read-only enumeration of the complete invocable surface \u2014 every built-in CLI command with aliases and hidden/gated status, every bundled skill, and every component of every installed plugin across all marketplaces; reads the shipped binary because upstream publishes no built-in command list, and carries an integrity verdict so a drifted build reports counts as floors rather than silently short totals), audit-install-state (read-only audit of the machine-scope ~/.claude installation directory and ~/.claude.json \u2014 full inventory split into an authored surface and rolled-up bulk trees, product-managed retention vs genuinely unmanaged state, filename-scheme resolution before any process-liveness check, and deliberate/mid-experiment detection; reports, never deletes), audit-performance (read-only slowness-diagnostic capture run at the moment the machine or a session feels slow: CLI version, retention-sweep health including the silent unparsable-settings pause, a timed census walk of the install tree as a sweep-cost proxy, active-session and plugin-fleet counts, a process census, and the fan-out layer, which covers a load-labelled no-op spawn baseline, every hook that will fire bucketed per-tool-call versus per-turn with its invocation shape, the configured statusline, subagent concurrency and spawn-depth ceilings against documented defaults, whether running sessions predate the settings file they are judged by, and orphan attribution by parent liveness rather than age; read against a bundled known-performance-issues reference that also records the causes tested and cleared; separates the four documented suspects of accumulated state, version regression, component bloat, and per-spawn fan-out cost, and routes remediation out; reports, never mutates, and never executes a discovered hook or statusline command), audit-native-overlap (map native Claude Code surfaces \u2014 built-in CLI commands, bundled skills, plugin-backed built-ins, session-provided skills \u2014 against the current repo's plugin skills and agents, so a custom component never silently duplicates what Claude Code itself ships; bare invocation is a read-only overlap report carrying the extraction's integrity floors and a shared-listing-budget exposure section, verdicts are human-gated in a committed store rendered into a generated registry whose every row carries an observable recheck trigger, and only an explicit apply step bakes presence-gated native references into descriptions and Boundary sections), observability (read locally captured telemetry \u2014 OTEL store, collector, hook-event JSONL, ccusage \u2014 with trend reports and store pruning), known-issues (search known Claude product GitHub bugs, check service health, maintain a persistent tracked-issue registry), changelog (ingest Claude Code changelog entries and integrate them into the current repo), plugins (bring a machine's plugin fleet current on demand \u2014 marketplace refresh, effective-scope updates including in-repo project/local installs, new-plugin install per policy, scope-divergence detection and explicit convergence), morning-brief (read-only gh-based operator morning view \u2014 queue-label counts, merge-ready PRs, parked decisions with their RECOMMENDED lines, and loop-lane telemetry freshness), lanes (start/restart/stop/status loop lanes as named background Claude Code sessions seeded from canonical prompt files, with per-lane model/effort, a repo-pull + marketplace-refresh launch step, and a consume-restarts action \u2014 an OS-schedulable reader that relaunches stopped lanes whose telemetry carries a restart_request), and a re-runnable setup action that settles where the known-issues registry lives. Plus a family of eight advisory *-audit hooks (API errors, config changes, instruction loads, permission denials, pre-compaction, skill usage, tool failures, and unsurfaced hook failures \u2014 the last also warns the user via systemMessage, since a hook that fails to launch enforces nothing and Claude Code surfaces the failure to nobody) that emit the shared hook-telemetry envelope, and a reference sink that maps envelopes into the hook-events.jsonl the observability skill reads.", "author": { "name": "Melodic Software", "email": "info@melodicsoftware.com" diff --git a/plugins/claude-ops/CHANGELOG.md b/plugins/claude-ops/CHANGELOG.md index 0df5a62692..8bdd473afd 100644 --- a/plugins/claude-ops/CHANGELOG.md +++ b/plugins/claude-ops/CHANGELOG.md @@ -3,6 +3,55 @@ All notable changes to the `claude-ops` plugin are documented here. Format follows [Keep a Changelog](https://keepachangelog.com/en/1.1.0/); this plugin uses semantic versioning. +## [0.39.0] + +### Fixed + +- **`audit-skill-visibility`'s starvation band now carries a usage signal.** + `compute_listing` sorted on a `usage_score` that only the test fixtures ever + set: the live collector builds its denominator from a filesystem walk, and + usage events were joined afterwards, so every real run scored zero and the band + fell through to its alphabetical tiebreaker while presenting itself as + usage-informed. On this machine that ranked `adhd:clarify` (1 use) as first to + lose its description and `work-items:triage` (99 uses) as among the safest. + Scores are now computed before the listing is built. +- **Truncation is modelled as the greedy first-fit walk the product runs, not a + score-ordered prefix.** The product's grant loop has no early exit, so it walks + every competing entry with a running description budget and a cheap low-scored + description can be granted after an expensive higher-scored one was refused. + Description length is therefore a second ranking input. The budget itself is + computed forward from a floor, `budget - V`, where V is what the listing costs + before any description is granted; deriving it from the overflow put the + grant boundary in the wrong place. Ties keep catalog order rather than + alphabetical, matching the product's stable sort, which decides the entire + ordering when nothing is scored. +- **Usage recorded under a skill's bare leaf no longer vanishes.** Events were + looked up by qualified `:` name only, while the stores hold both + that key and the bare leaf as separate rows, so the bare row was discarded with + nothing saying so. `source-control:babysit-prs` reported 97 invocations against + an actual 475. A bare key is now attributed when exactly one skill owns that + leaf, and withheld with its candidates when more than one does. + +### Added + +- **The listing scorer is mirrored rather than guessed at.** `listing_score` + implements `usageCount * max(0.5 ** (daysSinceUse / 7), 0.1)`, recovered from + Claude Code 2.1.251 and carrying a verification stamp with its basis and + recheck trigger. The ordering is decay-weighted, so "least invoked" was never + the right description of it: a heavily used but stale skill can rank below a + lightly used fresh one. Evidence in + `skills/audit-skill-visibility/reference/listing-scorer.md`, with the counters' own + semantics and their seeding and throttle traps in the companion + `reference/usage-counters.md`. +- **`listing.score_basis`, so an unscored band admits it.** When no usage + survives to weigh, the order is catalog order and nothing more; the basis reads + `unscored` and competing rows carry `confidence: "unscored"` rather than + borrowing `inferential`, which claims more than the data supports. The basis is + decided over the CONTENDERS, not the whole denominator: an exempt skill + carrying real usage must not label a contest whose every entrant is at zero. + The markdown report renders the distinction too, rather than describing an + unranked order as a likelihood band. + ## [0.38.23] ### Fixed diff --git a/plugins/claude-ops/skills/audit-skill-visibility/SKILL.md b/plugins/claude-ops/skills/audit-skill-visibility/SKILL.md index 2cddac0814..cbc16cba22 100644 --- a/plugins/claude-ops/skills/audit-skill-visibility/SKILL.md +++ b/plugins/claude-ops/skills/audit-skill-visibility/SKILL.md @@ -1,5 +1,5 @@ --- -description: "Audit whether each installed skill is actually VISIBLE to the model, and diagnose why most of a fleet never gets used. A skill is invisible when its description is dropped by the skill-listing context budget (Claude Code drops descriptions starting with the least-invoked skills, so an unused skill loses the keywords that would let it be matched and stays unused), when frontmatter is malformed or a description is missing, when skillOverrides or a disabled plugin hides it, or when disable-model-invocation keeps it out of context by design. Reports reachability, observed usage, and whether it is losing the budget contest. Computing whether the listing overflows from documented settings, and withholding every verdict the data cannot support rather than reporting absence of data as absence of use. Read-only; never disables, deletes, or edits a skill. Use when: 'why do I never use most of my skills', 'why does Claude never suggest this skill', 'are my skill descriptions being dropped', 'is my skill listing over budget', 'which skills can the model actually see', 'which skills are starved', 'I have too many skills to know when to use them', 'audit skill visibility'. Not for: which skills are unused versus their context cost as a one-shot check (Claude Code ships that in /doctor and the Stats tab), repo-authoring listing-budget lint (use skill-quality's check-listing-budget), enumerating what is installed (use /claude-ops:inventory), or reading telemetry infrastructure (use /claude-ops:observability)." +description: "Audit whether each installed skill is actually VISIBLE to the model, and diagnose why most of a fleet never gets used. A skill is invisible when its description is dropped by the skill-listing context budget (Claude Code drops descriptions by a decay-weighted usage score, so an unused skill loses the keywords that would let it be matched and stays unused), when frontmatter is malformed or a description is missing, when skillOverrides or a disabled plugin hides it, or when disable-model-invocation keeps it out of context by design. Reports reachability, observed usage, and whether it is losing the budget contest. Computing whether the listing overflows from documented settings, and withholding every verdict the data cannot support rather than reporting absence of data as absence of use. Read-only; never disables, deletes, or edits a skill. Use when: 'why do I never use most of my skills', 'why does Claude never suggest this skill', 'are my skill descriptions being dropped', 'is my skill listing over budget', 'which skills can the model actually see', 'which skills are starved', 'I have too many skills to know when to use them', 'audit skill visibility'. Not for: which skills are unused versus their context cost as a one-shot check (Claude Code ships that in /doctor and the Stats tab), repo-authoring listing-budget lint (use skill-quality's check-listing-budget), enumerating what is installed (use /claude-ops:inventory), or reading telemetry infrastructure (use /claude-ops:observability)." argument-hint: "[--installed [dir]] [--plugins-root ] [--render markdown|json] [--now ] [--fixture ]. Collects live; --installed reads the plugin manifest, else fleet defaults to ./plugins" user-invocable: true disable-model-invocation: false @@ -23,16 +23,20 @@ my skill fleet never get used?* A skill the model cannot see cannot be chosen, s Claude Code budgets the model-visible skill listing at a fraction of the context window (`skillListingBudgetFraction`, default 0.01) and, when it overflows, -**drops descriptions starting with the skills you invoke least**. Names always -survive, descriptions do not. A skill at zero usage therefore loses its +**sheds descriptions from the lowest-scoring skills first**. Names always +survive, descriptions do not. A skill at zero usage scores zero, so it loses its description, loses the keywords a request would match against, and stays at -zero. Unused is partly self-causing, and the loop is documented: the budget -fraction and per-entry cap are owned by +zero. Unused is partly self-causing. + +The budget fraction and per-entry cap are owned by (`skillListingBudgetFraction`, -`skillListingMaxDescChars`) and the drop behavior by - ("Skill descriptions are cut short"). -Verified 2026-08-31; recheck trigger: a fetch of either page no longer matching -this paragraph re-derives it and the scripts' `ListingConfig` defaults. +`skillListingMaxDescChars`); that page is authoritative and matches. **The drop +ORDER is not.** ("Skill descriptions are +cut short") says "starting with the skills you invoke least", still as of +2026-08-31; the binary ranks by a decay-weighted score and then walks the list +first-fit, so neither the ordering nor the guarantee holds. Take the ordering +from the binary: [reference/listing-scorer.md](reference/listing-scorer.md) +carries the counterexamples, the greps, and the stamp. So the useful question is not *which skills are unused*. Claude Code already reports that in `/doctor` and the Stats tab. It is **which skills are starved by @@ -168,6 +172,14 @@ a user as documented. - **Ambiguous attribution is reported, not guessed.** Two marketplaces shipping a same-named plugin collapse to one usage key; those rows are marked `ambiguous-attribution` rather than attributed to one of them. +- **Two possible usage keys per skill.** The stores hold the qualified + `:` key and the bare leaf as separate rows. Both are collected; a + bare key is attributed only when exactly one skill owns that leaf, ambiguous + ones are withheld with their candidates. +- **The starvation band is decay-weighted, not a count**, and reports + `score_basis: "unscored"` when nothing survives to weigh. Read + [reference/listing-scorer.md](reference/listing-scorer.md) before changing that + ordering or quoting it to a user. ## Scope boundary diff --git a/plugins/claude-ops/skills/audit-skill-visibility/reference/listing-scorer.md b/plugins/claude-ops/skills/audit-skill-visibility/reference/listing-scorer.md new file mode 100644 index 0000000000..fc1fd12cb5 --- /dev/null +++ b/plugins/claude-ops/skills/audit-skill-visibility/reference/listing-scorer.md @@ -0,0 +1,113 @@ +# The listing scorer, mirrored + +Read this when the starvation band's ordering is in question: why it is not a +count, why a bare usage key does not move it, and when it means nothing at all. +The engine's own stamp lives beside `listing_score` in +`scripts/audit_skill_visibility.py`; this file is the reasoning, not a second +copy of the claim. + +## What the product actually does + +Claude Code ranks skills for description truncation by + +```text +usageCount * max(0.5 ** (daysSinceUse / 7), 0.1) +``` + +then sorts that score descending and walks **every** competing entry with a +running description budget, granting whatever still fits and rendering the rest +name-only. + +**It is a greedy first-fit walk, not a score-ordered prefix.** The grant loop has +no early exit, so a cheap low-scored description can still be granted after an +expensive higher-scored one was refused. Description LENGTH is therefore a second +ranking input, which no prose account of this mechanism mentions: + +```js +for (let me of W) { + let ge = me.entryLen - (me.cmd.name.length + 2); + if (ge <= pe) pe -= ge; else fe.push(me); // no break +} +``` + +This matters for the report's two fields. The `verdict` mirrors the walk. The +`band` ranks exposure, lowest score first, and the two are allowed to disagree: +a band-1 row with a very short description can survive a pass that sheds a +better-scored row with a long one. + +Recovered from `claude.exe` at Claude Code 2.1.251 (`zPe` the scorer, `Ymt` the +truncator), then re-verified unchanged at 2.1.252. The scorer has two further +call sites, the slash-menu top-five pin and the command-search score boost, which +is corroboration that it is the product's general usage-priority function rather +than a listing-local helper. What the counts it reads actually mean, including +the seeding and throttle traps, is [usage-counters.md](usage-counters.md). + +## Why "least invoked" was the wrong description + +The score is decay-weighted with a seven-day half life and a floor at a tenth, +so recency competes with volume. A skill used 100 times sixty days ago scores +`100 * 0.1 = 10` and loses its description to one used 12 times today, which +scores 12. Any wording that says descriptions are shed "starting with the +least-invoked skills" describes a mechanism the product does not have. + +The floor matters at both ends. A never-used skill scores exactly zero and sorts +last, so it loses its description first under any material overflow, which is the +feedback loop this whole skill exists to expose. It is not *always* shed, though: +because the walk is first-fit, a zero-scored skill with a very short description +can still be granted from what the others left. A once-used skill never decays +below `0.1 * usageCount`, so it never falls back into the never-used band. + +Two more properties the ordering depends on: + +- **Ties keep catalog order, not alphabetical.** The product's sort is stable, so + equal scores stay in input order. This decides everything in the `unscored` + case, where every score is zero and the tiebreak IS the whole ordering. +- **The grant budget is computed forward from a floor**, `budget - V`, where `V` + is what the listing costs before any description is granted: every listed entry + pays for its own name, the exempt classes pay their full rendering, and the + separators are charged too. Deriving it by subtracting the overflow instead + gets the grant boundary wrong, not just the ordering. + +## Why a bare usage key does not move the band + +The stores record a skill's usage under either its qualified `:` +name or its bare leaf, as separate rows. `zPe` looks up the listing entry's name +directly and does no fallback between the two; the product's own display helper +(`oKn`) does, but the scorer does not. + +So the mirror does not either. Feeding it the merged total would predict a +truncation the product will not perform, and predicting the product wrongly is +the one thing this band must not do. `observation` asks a different question, +"how much has this skill been used", and takes every event, merged. + +That means the product can leave an entry unscored while the skill is heavily +used under its other key. Faithfully reproducing that is the point. + +## When the band means nothing + +`listing.score_basis` is `unscored` whenever no usage survives to weigh. The +ordering is then the alphabetical tiebreaker and carries no signal, so competing +rows report `confidence: "unscored"` rather than `inferential`. The distinction +is load-bearing: `inferential` claims a ranking exists and may be imprecise; +`unscored` says no ranking was possible. + +## On drift + +This is a minified bundle, not a published interface. The recheck trigger is a +release note naming the skill listing, its character budget, or the usage +counters, or the counters changing shape in `~/.claude.json`. On a mismatch the +honest degradation is back to `unscored`, never a confidently wrong band. + +**Locate the scorer by shape, never by name.** The minified identifier moves +between builds. It was `zPe` in 2.1.251 and `WPe` in 2.1.252, with a +byte-identical body, so a recheck that greps the old name finds nothing and +concludes the mechanism was removed. Grep the arithmetic instead: + +```bash +grep -a -o -E '.{0,180}Math\.pow\(0\.5,.{0,180}' "$(command -v claude)" +grep -a -o -E '.{0,260}budgetTruncatedSkills:.{0,60}' "$(command -v claude)" +``` + +The first run of that recheck happened the same day the stamp was written: the +CLI auto-updated from 2.1.251 to 2.1.252 mid-session, the trigger fired, and both +the formula and the descending-sort truncation came back unchanged. diff --git a/plugins/claude-ops/skills/audit-skill-visibility/reference/usage-counters.md b/plugins/claude-ops/skills/audit-skill-visibility/reference/usage-counters.md new file mode 100644 index 0000000000..a2b647feb6 --- /dev/null +++ b/plugins/claude-ops/skills/audit-skill-visibility/reference/usage-counters.md @@ -0,0 +1,130 @@ +# The `~/.claude.json` usage counters + +Read this when a count this skill reports looks wrong, or before writing a new +consumer of these counters. It records what each region actually means, and the +four traps that make a naive read wrong. `~/.claude.json` is the user-scope state +file, a different file from `~/.claude/settings.json`. + +Companion: [listing-scorer.md](listing-scorer.md), which covers how the product +ranks these counts when it truncates the skill listing. + +## Verification stamp + +Follows `docs/conventions/upstream-drift`. + +- **Claim:** the write paths, seeding behavior, and per-region semantics below. +- **Basis:** string extraction of `claude.exe` at Claude Code 2.1.251, the write + and read functions named per region, confirmed against a live `~/.claude.json`. + Scorer and truncator re-verified unchanged at 2.1.252. +- **As-of:** 2026-08-31. +- **Locate by shape, never by name.** Minified identifiers move between builds; + the listing scorer was `zPe` at 2.1.251 and `WPe` at 2.1.252 with an identical + body. Grep the arithmetic and the field names, not the function names. +- **Recheck trigger:** a release note naming skill usage counters, the skill + listing, or `/doctor`'s unused-component check; or these keys gaining or losing + a field. + +## `skillUsage`, machine-global, per skill + +`{"": {"usageCount": , "lastUsedAt": }}` + +Write path, reduced to its decisive lines: + +```js +let d = o.skillUsageLastWriteAt.get(t); +if (d !== void 0 && u - d < 60000) return; +o.skillUsageLastWriteAt.set(t, u), + Ae((y) => ({ ...y, skillUsage: { ...y.skillUsage, + [t]: { usageCount: (k?.usageCount ?? 0) + 1, lastUsedAt: u } } }), r) +``` + +- Written on real skill dispatch only. **No install-time or session-start + seeding**, so a skill's `lastUsedAt` is trustworthy evidence of actual use. + This is the one place it is trustworthy; see `pluginUsage` below. +- `usageCount` is a lifetime total since install. It never resets and is never + windowed, which is why the engine's `T-baseline` tier supports a lifetime + claim and refuses a windowed one. +- **The 60-second per-skill throttle DROPS the increment, it does not coalesce + it.** A skill invoked five times in one minute records one. The undercount is + unbounded and unrecoverable, so for loop-driven or rapid-fire skills this + counter is a floor rather than a count. Divergence from OTEL is expected here + and must not be reconciled away. +- **Keys are inconsistent between qualified and bare form.** One machine held + both `babysit-prs` (378) and `source-control:babysit-prs` (97) as separate rows + for the same skill. Every consumer has to resolve both spellings. Note that the + product's own listing scorer does NOT do this fallback, which is why the + starvation band deliberately does not either. + +## `pluginUsage`, machine-global, per plugin + +`{"@": {"usageCount": , "lastUsedAt": , "lastUsedNumStartups": }}` + +Three write paths, and the second is the trap: + +- Batched flush accumulates: `usageCount: (d?.usageCount ?? 0) + u.count`. Unlike + `skillUsage`, nothing is lost to throttling. +- Absent entries are **seeded** with `usageCount: 0`, `lastUsedAt: now`, + `lastUsedNumStartups: ` on install or enable. +- `lastUsedAt` and `lastUsedNumStartups` are **refreshed on re-enable**, with no + usage at all. + +So for a plugin, `lastUsedAt` is usage evidence only when `usageCount > 0`. A +zero-count plugin's `lastUsedAt` is the seed time and means nothing. Measured on +one machine: 46 of 65 plugins looked "used today" while none had been used. That +measurement is what `is_usage_evidence()` in the engine exists to enforce. + +**A plugin "use" is broad and not comparable across plugin shapes.** Usage is +recorded whenever a slash command, skill, agent, MCP tool or resource, or hook is +dispatched from that plugin, plus LSP servers delivering diagnostics or code +navigation. Hook-only plugins therefore dominate any raw ranking: on one machine +`guardrails` showed 200,804 and `context-guard` 106,752, against low hundreds for +skill-shaped plugins. A hook plugin's count is a per-tool-call tally; a skill +plugin's count is an invocation tally. Ranking plugins by raw count ranks them by +hook chattiness, not by value. + +`pluginUsageLspGraceAppliedIds` at the top level records which plugin ids got the +LSP grace backfill, so a lifetime zero on an LSP-only plugin may simply predate +the tracking. + +## `agentLastUsed` + +Observed holding a single key (`bg`). Subagent usage is effectively **not tracked +here**, so no agent-usage analysis can be sourced from this file. The transcript +parsers in the session-flow retro skills are the working source for that. + +## `projects[]`, per-project, last session only + +Each entry is a **snapshot of the last session in that directory, not a running +total**. The object is rebuilt from live session getters at each write, and the +matching exit telemetry reports the same values as `last_session_*`. Every +session end overwrites the previous one. + +It carries cost, wall and API durations, lines added and removed, the four token +totals, web-search requests, session id and start time, frame-duration +percentiles, and `lastModelUsage` broken down per model id. + +**The decisive fact for per-project usage questions:** `skillUsage` and +`pluginUsage` are top level ONLY. Nothing under `projects` carries them, so this +file structurally cannot answer "which skills does this project use". That +question needs the plugin's own `skill-usage.jsonl` store, whose rows carry +`project`, `project_id`, and `branch`, or OTEL. + +## Growth, and the one supported shrink lever + +The file is never swept: `cleanupPeriodDays` does not reach it, because it lives +in the home directory rather than under `~/.claude`. The supported lever is +`claude project purge `, which removes one project's entry. Every running +session polls the file at 1 Hz, and the strongest public report of curing input +lag pruned this file rather than the install tree. + +`.claude.json.tmp..` siblings are failed atomic-write remnants; the +leading number only looks like a PID. See +[`audit-install-state/reference/surfaces.md`](../../audit-install-state/reference/surfaces.md) +for the install-tree side of this, which holds the file stat-only because its +values can carry tokens. + +## Incidental counters, not usage signals + +`numStartups`, `promptQueueUseCount`, `tipsHistory`, `tipLifetimeShownCounts`, +`passesUpsellSeenCount`, `lspRecommendationIgnoredCount`. These drive tip +cooldowns and onboarding state, not component usage analysis. diff --git a/plugins/claude-ops/skills/audit-skill-visibility/scripts/audit_skill_visibility.py b/plugins/claude-ops/skills/audit-skill-visibility/scripts/audit_skill_visibility.py index 27557f8bb6..35c89ee728 100755 --- a/plugins/claude-ops/skills/audit-skill-visibility/scripts/audit_skill_visibility.py +++ b/plugins/claude-ops/skills/audit-skill-visibility/scripts/audit_skill_visibility.py @@ -106,6 +106,95 @@ def tier_supports(tier: str, claim: str) -> bool: return claim in TIER_CAPABILITIES.get(tier, set()) +# Claude Code's own listing-budget scorer, mirrored so the starvation band can +# predict which descriptions the product will actually drop. Before this existed +# the band sorted on a field nothing populated, so it rendered alphabetical order +# dressed as usage-informed. +# +# -- Verification stamp (docs/conventions/upstream-drift) ---------------------- +# Claim: the product ranks skills for description truncation by +# `usageCount * max(0.5 ** (daysSinceUse / 7), 0.1)`, sorts that score +# descending, then walks EVERY competing entry with a running description +# budget, granting whatever fits and rendering the rest name-only. The walk is +# greedy first-fit with no early exit, so it is not a score-ordered prefix and +# description length is a second ranking input. +# Basis: string extraction of `claude.exe`, first at Claude Code 2.1.251 (`zPe` +# the scorer, `Ymt` the truncator, plus the scorer's two other call sites, the +# slash-menu top-5 pin and the command-search score boost), then RE-VERIFIED +# unchanged at 2.1.252. Evidence in +# reference/listing-scorer.md, beside this skill. +# As-of: 2026-08-31, re-verified against Claude Code 2.1.252. +# Locate it by SHAPE, never by name. The minified identifier is not stable across +# builds: the scorer was `zPe` in 2.1.251 and `WPe` in 2.1.252, with a +# byte-identical body. Grep for the arithmetic instead, e.g. +# `grep -a -o -E '.{0,180}Math\.pow\(0\.5,.{0,180}' `, and for +# the truncator `budgetTruncatedSkills`. +# Reasoning and the counters' own semantics: reference/listing-scorer.md and +# reference/usage-counters.md, beside this skill. +# Recheck trigger: a release note naming the skill listing, its character budget, +# or skill usage counters; or the counters changing shape in `~/.claude.json`. +# On mismatch: the report must degrade to `score_basis: "unscored"` and say the +# ordering is unknown. A confidently wrong band is worse than no band. +# ----------------------------------------------------------------------------- +LISTING_SCORE_HALF_LIFE_DAYS = 7.0 +LISTING_SCORE_FLOOR = 0.1 + + +def listing_score(count: int, last_used: datetime | None, clock: datetime) -> float: + """Mirror of the product's scorer. Zero when there is no usage to weigh.""" + if count <= 0 or last_used is None: + return 0.0 + days = (clock - last_used).total_seconds() / 86400.0 + decay = 0.5 ** (days / LISTING_SCORE_HALF_LIFE_DAYS) + return count * max(decay, LISTING_SCORE_FLOOR) + + +def resolve_event_keys( + denominator: list[dict], events: list[dict] +) -> tuple[dict[str, list[dict]], list[tuple[str, list[str]]]]: + """Group events by the qualified skill they belong to. + + The stores record a skill's usage under either its qualified `:` + name or its bare leaf, and the two land as separate rows. This machine holds + both `babysit-prs` and `source-control:babysit-prs`. Looking events up by + qualified name alone silently discarded every bare-key row, which is a + reported count that is simply too low with nothing saying so. + + A bare key is attributed only when exactly one skill in the fleet carries + that leaf. Piling an ambiguous leaf onto one plugin would invent usage, the + same failure the `custom_skill` redaction guard exists to prevent, so an + ambiguous key is returned for the withheld section instead of being spent. + """ + qualified = {entry["qualified_name"] for entry in denominator} + # Owners are counted per ENTRY, not per distinct qualified name. Two + # marketplaces shipping the same plugin produce two denominator rows with an + # identical qualified name, which the report already marks + # `ambiguous-attribution`. Collapsing them into a set would let a bare key + # pass the single-owner test and then be reported on BOTH rows, inventing + # usage for an attribution the audit already knows it cannot make. + leaf_owners: dict[str, list[str]] = defaultdict(list) + for entry in denominator: + name = entry["qualified_name"] + leaf = name.split(":", 1)[1] if ":" in name else name + leaf_owners[leaf].append(name) + + by_skill: dict[str, list[dict]] = defaultdict(list) + ambiguous: dict[str, list[str]] = {} + for event in events: + key = event.get("skill") + if key is None: + continue + if key in qualified: + by_skill[key].append(event) + continue + owners = leaf_owners.get(key, []) + if len(owners) == 1: + by_skill[owners[0]].append(event) + elif len(owners) > 1: + ambiguous[key] = sorted(set(owners)) + return by_skill, sorted(ambiguous.items()) + + def parse_native( skill_usage: dict, first_start: datetime ) -> tuple[list[dict], datetime]: @@ -708,56 +797,127 @@ def _demand_chars(entry: dict, cfg: ListingConfig) -> int: return min(len(description) + joiner + len(when_to_use), cfg.max_desc_chars) -def compute_listing(denominator: list[dict], cfg: ListingConfig) -> dict: +def compute_listing( + denominator: list[dict], + cfg: ListingConfig, + scores: dict[str, float] | None = None, +) -> dict: """Budget arithmetic, split by confidence. CERTAIN: whether the listing overflows and by how much -- pure arithmetic over documented settings against summed description lengths. INFERENTIAL: which particular skills lose their descriptions. That ordering - comes from an undocumented scorer pinned to one build, so it is rendered as - a ranked band and labelled, never as an exact cutoff. + comes from a scorer recovered from one build of the product (see + `listing_score`), so it is rendered as a ranked band and labelled, never as + an exact cutoff. + + `scores` carries the mirrored scorer's output per qualified name. When it is + absent or empty the ordering has no usage signal behind it, and the returned + `score_basis` says so rather than letting the alphabetical tiebreaker pass + for a usage ranking. """ budget = listing_budget_chars(cfg) + scores = scores or {} rows: list[dict] = [] demand = 0 + # The grant loop's budget is computed FORWARD from a floor, never backward + # from the overflow: the product starts at `budget - V`, where V is what the + # listing costs before any description is granted. Every entry that appears + # at all pays for its own name; the exempt classes pay their full rendering + # because they are never candidates. + floor = 0 + listed = 0 for entry in denominator: eligibility = _eligibility(entry) - chars = _demand_chars(entry, cfg) if eligibility == "competing" else 0 + desc_chars = _demand_chars(entry, cfg) + chars = desc_chars if eligibility == "competing" else 0 demand += chars + # A `disable-model-invocation` skill is absent from the listing + # ENTIRELY, so unlike the other two exempt classes it costs nothing and + # takes no separator. + if eligibility != "exempt-user-only": + listed += 1 + name_chars = len(entry["qualified_name"]) + if eligibility == "exempt-bundled": + # Keeps its description unconditionally, so it is charged for it. + floor += name_chars + 4 + desc_chars + else: + floor += name_chars + 2 rows.append( { "qualified_name": entry["qualified_name"], "eligibility": eligibility, "demand_chars": chars, - "usage_score": entry.get("usage_score", 0), + "usage_score": scores.get( + entry["qualified_name"], entry.get("usage_score", 0) + ), } ) overflow = max(0, demand - budget) verdict = "overflowing" if overflow > 0 else "listing-fits" + floor += max(0, listed - 1) + competing = [r for r in rows if r["eligibility"] == "competing"] - # Least-used first: that is the order the product drops descriptions in, so - # rank 1 is the most likely to have already lost its description. - competing.sort(key=lambda r: (r["usage_score"], r["qualified_name"])) - # Descriptions are dropped least-invoked-first only UNTIL the listing fits, - # so the starved set is the prefix whose demand covers the overflow -- not - # the whole fleet. Marking every competing row starved on a one-character - # overflow would libel exactly the skills the mechanism protects longest, - # and the renderer would tell the user they are running name-only. - remaining = overflow - for rank, row in enumerate(competing, start=1): - row["band"] = rank if overflow > 0 else None - row["confidence"] = "inferential" if overflow > 0 else "certain" + + # Basis is decided by whether any score survived AMONG THE CONTENDERS, not + # by whether a scores argument arrived and not over the whole denominator. A + # bundled, name-only, or disable-model-invocation skill can carry real native + # usage while being excluded from the contest entirely; counting its score + # here would label a listing `native-counters` whose every actual contender + # is at zero, so a pure catalog ordering would be dressed as `inferential`. + # That is the exact defect this report exists to stop, one scope up. + score_basis = ( + "native-counters" if any(r["usage_score"] for r in competing) else "unscored" + ) + + # WHICH rows are shed is a GREEDY FIRST-FIT walk, not a score-ordered + # prefix. The product sorts descending by score and then walks EVERY + # competing entry with a running description budget, granting whatever still + # fits and shedding whatever does not. Crucially its loop has no early exit, + # so a cheap low-scored description can still be granted after an expensive + # higher-scored one was refused. Modelling this as a prefix understated the + # protection long descriptions lose and overstated it for short ones. + # + # Consequence worth stating plainly: description LENGTH is a ranking input, + # which no prose description of this mechanism mentions. + # + # Ties keep CATALOG ORDER, not alphabetical. The product's sort is stable, so + # equal scores stay in input order; Python's is too, which is why this sorts + # on the score alone. It matters most in the `unscored` case, where every + # score is 0 and the tiebreaker IS the whole ordering: an alphabetical one + # would disagree with the product on every row. + competing.sort(key=lambda r: -r["usage_score"]) + remaining = max(0, budget - floor) + for row in competing: if overflow <= 0: row["verdict"] = "listing-fits" - elif remaining > 0: - row["verdict"] = "likely-starved" + elif row["demand_chars"] <= remaining: remaining -= row["demand_chars"] - else: row["verdict"] = "likely-retained" + else: + row["verdict"] = "likely-starved" + + # The band is a separate question from the verdict: it ranks how exposed a + # row is, lowest score first, so band 1 is the row the mechanism protects + # least. It can disagree with the verdict, and that disagreement is real + # rather than a bug -- a band-1 row with a very short description can survive + # a first-fit pass that sheds a better-scored row with a long one. + by_exposure = sorted(competing, key=lambda r: r["usage_score"]) + for rank, row in enumerate(by_exposure, start=1): + row["band"] = rank if overflow > 0 else None + # An unscored ordering is alphabetical, so its band carries no signal at + # all. That is a weaker claim than an inferential one and must not wear + # the same label. + if overflow <= 0: + row["confidence"] = "certain" + elif score_basis == "unscored": + row["confidence"] = "unscored" + else: + row["confidence"] = "inferential" for row in rows: if row["eligibility"] != "competing": row["band"] = None @@ -772,6 +932,7 @@ def compute_listing(denominator: list[dict], cfg: ListingConfig) -> dict: "demand_chars": demand, "overflow_chars": overflow, "verdict": verdict, + "score_basis": score_basis, "competing_count": len(competing), "starved_count": sum(1 for r in competing if r["verdict"] == "likely-starved"), "exempt_count": len(rows) - len(competing), @@ -902,7 +1063,34 @@ def classify( ) -> dict: """Pure. Fleet + events + config + clock + horizons -> report model.""" tier = resolve_tier(set(horizons)) - listing = compute_listing(denominator, listing_config or ListingConfig()) + events_by_skill, ambiguous_keys = resolve_event_keys(denominator, events) + + # The band mirrors the product, so it is scored the way the product scores: + # from the NATIVE counters only, under the EXACT qualified key. `zPe` does no + # bare-key fallback of its own -- when usage was recorded under a bare leaf + # the product's own scorer sees zero for that listing entry too, so scoring + # the merged total here would predict a truncation the product will not + # perform. The merge below is for the observation count, which asks a + # different question and wants every event. + native_scores: dict[str, float] = {} + for entry in denominator: + name = entry["qualified_name"] + native = [ + e + for e in events_by_skill.get(name, []) + if e.get("source") == "native" and e.get("skill") == name + ] + if not native: + continue + native_scores[name] = listing_score( + sum(int(e.get("count", 1)) for e in native), + max(e["ts"] for e in native), + clock, + ) + + listing = compute_listing( + denominator, listing_config or ListingConfig(), native_scores + ) starvation_by_name = {r["qualified_name"]: r for r in listing["skills"]} # Narrowest horizon = the most recent start = the least we can see back to. # Two different questions, two different horizons. @@ -923,13 +1111,22 @@ def classify( for entry in denominator: seen[entry["qualified_name"]] += 1 - events_by_skill: dict[str, list[dict]] = defaultdict(list) - for event in events: - events_by_skill[event["skill"]].append(event) - skills: list[dict] = [] withheld: list[dict] = [] + for key, owners in ambiguous_keys: + withheld.append( + { + "skill": key, + "claim": "observation", + "reason": ( + f"bare usage key `{key}` matches {len(owners)} skills " + f"({', '.join(owners)}); attributing it would invent usage " + "for whichever one was picked" + ), + } + ) + for entry in denominator: name = entry["qualified_name"] own_events = events_by_skill.get(name, []) @@ -1120,19 +1317,38 @@ def _render_markdown(model: dict) -> str: f"{listing['overflow_chars']:,} characters.** " f"{listing['competing_count']} skills compete for " f"{listing['budget_chars']:,} characters of description budget, " - f"and descriptions are dropped least-invoked-first — so roughly " + f"and descriptions are shed lowest-score-first, so roughly " f"**{listing['starved_count']}** of them are running name-only, " f"which is why the model stops matching requests to those.", "", f"The other {listing['competing_count'] - listing['starved_count']} " - f"competing skills keep their descriptions: only enough of the " - f"least-used tail is dropped to close the gap, not the whole set.", - "", - "*Which* particular skills lost theirs is inferential — the ordering", - "comes from an undocumented scorer. Treat the ranking as a likelihood", - "band, not a cutoff line.", + f"competing skills keep their descriptions. The score is " + f"decay-weighted, not a raw invocation count, and the walk grants " + f"whatever still fits rather than shedding a clean tail, so " + f"description length matters too.", "", ] + # An unscored run has no usage behind its ordering at all. Saying + # "inferential" there would repeat the exact defect this report + # exists to expose, one level up, so the two cases get different + # prose rather than a shared hedge. + if listing.get("score_basis") == "unscored": + lines += [ + "**No usage signal was available for any competing skill, so " + "*which* particular skills lost their descriptions is NOT " + "ranked here.** The order below is the catalog order and " + "carries no information about starvation likelihood. The " + "over-budget figure above is unaffected and still holds.", + "", + ] + else: + lines += [ + "*Which* particular skills lost theirs is inferential: the " + "ordering mirrors a scorer recovered from the shipped binary, " + "not a documented interface. Treat the ranking as a likelihood " + "band, not a cutoff line.", + "", + ] else: lines += [ f"Listing fits: {listing['demand_chars']:,} of " diff --git a/plugins/claude-ops/skills/audit-skill-visibility/scripts/test_audit_skill_visibility.py b/plugins/claude-ops/skills/audit-skill-visibility/scripts/test_audit_skill_visibility.py index 5be21fd6c3..3de07dd157 100755 --- a/plugins/claude-ops/skills/audit-skill-visibility/scripts/test_audit_skill_visibility.py +++ b/plugins/claude-ops/skills/audit-skill-visibility/scripts/test_audit_skill_visibility.py @@ -472,6 +472,215 @@ def test_no_band_when_the_listing_fits(self): self.assertIsNone(listing["skills"][0]["band"]) +class ListingScoreTest(unittest.TestCase): + """The mirrored scorer, and the refusal to dress zero up as a ranking.""" + + def test_score_decays_with_a_seven_day_half_life(self): + now = _utc(2026, 8, 31) + fresh = engine.listing_score(100, now, now) + one_half_life = engine.listing_score(100, now - timedelta(days=7), now) + self.assertAlmostEqual(fresh, 100.0) + self.assertAlmostEqual(one_half_life, 50.0) + + def test_decay_floors_at_a_tenth(self): + now = _utc(2026, 8, 31) + ancient = engine.listing_score(100, now - timedelta(days=3650), now) + self.assertAlmostEqual(ancient, 10.0) + + def test_a_stale_heavy_user_sorts_below_a_fresh_light_one(self): + """The whole reason `least invoked` was the wrong description.""" + now = _utc(2026, 8, 31) + stale = engine.listing_score(100, now - timedelta(days=60), now) + fresh = engine.listing_score(12, now, now) + self.assertLess(stale, fresh) + + def test_never_used_scores_zero(self): + now = _utc(2026, 8, 31) + self.assertEqual(engine.listing_score(0, now, now), 0.0) + self.assertEqual(engine.listing_score(5, None, now), 0.0) + + def test_all_zero_scores_report_an_unscored_basis(self): + """An alphabetical order must not be labelled a usage ranking.""" + entries = [ + { + "qualified_name": f"a:{i}", + "frontmatter": {"description": "x" * 1000}, + "plugin_enabled": True, + } + for i in range(10) + ] + listing = engine.compute_listing( + entries, engine.ListingConfig(context_window_tokens=200_000) + ) + self.assertEqual(listing["score_basis"], "unscored") + competing = [s for s in listing["skills"] if s["eligibility"] == "competing"] + self.assertTrue(all(s["confidence"] == "unscored" for s in competing)) + + def test_classify_scores_the_band_from_native_counters(self): + """Regression: the band used to sort on a field nothing populated.""" + now = _utc(2026, 8, 31) + entries = [ + { + "qualified_name": f"a:{i}", + "frontmatter": {"description": "x" * 1000}, + "plugin_enabled": True, + } + for i in range(10) + ] + # Counts ascend with the index, so the band must descend with it. + events = [ + {"skill": f"a:{i}", "ts": now, "source": "native", "count": (i + 1) * 10} + for i in range(10) + ] + model = engine.classify( + denominator=entries, + events=events, + config=engine.Config(), + clock=now, + horizons={"native": now - timedelta(days=400)}, + listing_config=engine.ListingConfig(context_window_tokens=200_000), + ) + self.assertEqual(model["listing"]["score_basis"], "native-counters") + bands = { + row["qualified_name"]: row["starvation"]["band"] for row in model["skills"] + } + # Least-scored is band 1, i.e. first to lose its description. + self.assertEqual(bands["a:0"], 1) + self.assertEqual(bands["a:9"], 10) + + +class ScoreBasisScopeTest(unittest.TestCase): + """The basis is decided by the contenders, not by the whole denominator.""" + + def test_usage_on_an_exempt_skill_does_not_score_the_contest(self): + """Regression: an exempt row's score used to flip the basis. + + A bundled, name-only, or user-only skill can carry real native usage + while being excluded from the contest entirely. Counting it labelled the + listing `native-counters` while every actual contender sat at zero, so a + pure catalog ordering got dressed as `inferential`. That is the defect + this whole report exists to expose, one scope up. + """ + entries = [ + { + "qualified_name": f"a:{i}", + "frontmatter": {"description": "x" * 1000}, + "plugin_enabled": True, + } + for i in range(10) + ] + entries.append( + { + "qualified_name": "a:manual", + "frontmatter": { + "description": "x" * 1000, + "disable_model_invocation": True, + }, + "plugin_enabled": True, + } + ) + listing = engine.compute_listing( + entries, + engine.ListingConfig(context_window_tokens=200_000), + # Only the exempt row has any usage at all. + {"a:manual": 500.0}, + ) + self.assertEqual(listing["score_basis"], "unscored") + competing = [s for s in listing["skills"] if s["eligibility"] == "competing"] + self.assertTrue(all(s["confidence"] == "unscored" for s in competing)) + + +class BareUsageKeyTest(unittest.TestCase): + """Usage recorded under a bare leaf must reach its qualified skill.""" + + def test_bare_key_is_attributed_when_the_leaf_is_unique(self): + now = _utc(2026, 8, 31) + model = engine.classify( + denominator=[_skill("source-control:babysit-prs")], + events=[ + {"skill": "babysit-prs", "ts": now, "source": "native", "count": 378}, + { + "skill": "source-control:babysit-prs", + "ts": now, + "source": "native", + "count": 97, + }, + ], + config=engine.Config(), + clock=now, + horizons={"native": now - timedelta(days=400)}, + ) + row = model["skills"][0] + self.assertEqual(row["observation"]["count"], 475) + + def test_ambiguous_bare_key_is_withheld_not_guessed(self): + now = _utc(2026, 8, 31) + model = engine.classify( + denominator=[_skill("toolchain:check"), _skill("skill-quality:check")], + events=[{"skill": "check", "ts": now, "source": "native", "count": 40}], + config=engine.Config(), + clock=now, + horizons={"native": now - timedelta(days=400)}, + ) + for row in model["skills"]: + self.assertEqual(row["observation"]["count"], 0) + withheld = [w for w in model["withheld"] if w["skill"] == "check"] + self.assertEqual(len(withheld), 1) + self.assertIn("toolchain:check", withheld[0]["reason"]) + self.assertIn("skill-quality:check", withheld[0]["reason"]) + + def test_a_bare_key_is_withheld_when_the_leaf_has_duplicate_entries(self): + """Two marketplaces shipping one plugin give two rows, one name. + + Collapsing owners into a set let the bare key pass the single-owner test + and then be reported on BOTH rows, inventing usage for an attribution the + report already marks `ambiguous-attribution`. + """ + now = _utc(2026, 8, 31) + model = engine.classify( + denominator=[_skill("dup:check"), _skill("dup:check")], + events=[{"skill": "check", "ts": now, "source": "native", "count": 40}], + config=engine.Config(), + clock=now, + horizons={"native": now - timedelta(days=400)}, + ) + for row in model["skills"]: + self.assertEqual(row["observation"]["count"], 0) + withheld = [w for w in model["withheld"] if w["skill"] == "check"] + self.assertEqual(len(withheld), 1) + self.assertIn("dup:check", withheld[0]["reason"]) + + def test_a_bare_key_does_not_score_the_band(self): + """`zPe` has no bare-key fallback, so the mirror must not add one.""" + now = _utc(2026, 8, 31) + entries = [ + { + "qualified_name": "a:one", + "frontmatter": {"description": "x" * 1000}, + "plugin_enabled": True, + }, + { + "qualified_name": "b:two", + "frontmatter": {"description": "x" * 1000}, + "plugin_enabled": True, + }, + ] + model = engine.classify( + denominator=entries, + events=[{"skill": "one", "ts": now, "source": "native", "count": 900}], + config=engine.Config(), + clock=now, + horizons={"native": now - timedelta(days=400)}, + listing_config=engine.ListingConfig(context_window_tokens=200_000), + ) + rows = {r["qualified_name"]: r for r in model["skills"]} + # The count reaches the observation field ... + self.assertEqual(rows["a:one"]["observation"]["count"], 900) + # ... but not the band, because the product's own scorer misses it too. + self.assertEqual(rows["a:one"]["starvation"]["usage_score"], 0) + self.assertEqual(model["listing"]["score_basis"], "unscored") + + class TierResolutionTest(unittest.TestCase): """A claim renders only at a tier that supports it.""" @@ -763,19 +972,80 @@ def test_one_char_overflow_starves_only_the_least_used_row(self): self.assertEqual(starved[0]["qualified_name"], "a:0") self.assertEqual(len(retained), 7) - def test_starved_set_covers_the_overflow_and_no_more(self): - # 10 x 1000 = 10_000 against 8_000 -> overflow 2_000 -> exactly 2 rows. + def test_the_name_floor_is_charged_before_any_description_is_granted(self): + """The grant budget starts at `budget - V`, not at the whole budget. + + Naive arithmetic says 10 x 1000 against 8_000 sheds exactly two rows. The + product charges every listed entry for its own name and the separators + first, so slightly less than the full budget is available to + descriptions, and a third row goes. Deriving the grant budget by + subtracting the overflow instead would miss that row entirely. + """ listing = self._listing(n=10, chars=1000, budget_tokens=200_000) + # The overflow figure itself is unchanged: it answers "does the listing + # overflow", which is a different question from "which entries win". self.assertEqual(listing["overflow_chars"], 2_000) starved = [s for s in listing["skills"] if s["verdict"] == "likely-starved"] - self.assertEqual(len(starved), 2) - self.assertEqual(sorted(s["qualified_name"] for s in starved), ["a:0", "a:1"]) + self.assertEqual(len(starved), 3) + # The three lowest-scored, since nothing here varies in length. + self.assertEqual( + sorted(s["qualified_name"] for s in starved), ["a:0", "a:1", "a:2"] + ) def test_most_used_skill_is_never_starved_while_others_can_absorb_it(self): listing = self._listing(n=10, chars=1000, budget_tokens=200_000) hottest = max(listing["skills"], key=lambda s: s["usage_score"]) self.assertEqual(hottest["verdict"], "likely-retained") + def test_a_cheap_low_scored_row_is_granted_after_a_costly_higher_one_is_shed(self): + """First-fit, not a prefix: the product's grant loop has no early exit. + + A prefix model sheds a contiguous run of the lowest-scored rows. The real + loop walks every entry, so a short description can still be granted after + a longer, better-scored one was refused. Description LENGTH is a ranking + input, which no prose account of this mechanism mentions. + """ + # Per-entry demand is capped at max_desc_chars (1536), so the budget is + # exhausted with full-cap fillers rather than one giant description. + entries = [ + { + "qualified_name": f"a:fill{i}", + "frontmatter": {"description": "x" * 1536}, + "plugin_enabled": True, + } + for i in range(5) + ] + entries += [ + # Outscores the cheap row, but cannot fit in what the fillers left. + { + "qualified_name": "a:hog", + "frontmatter": {"description": "x" * 1536}, + "plugin_enabled": True, + }, + # Lowest score in the set, but cheap enough to survive the leftovers. + { + "qualified_name": "a:cheap", + "frontmatter": {"description": "x" * 200}, + "plugin_enabled": True, + }, + ] + scores = {f"a:fill{i}": 100.0 - i for i in range(5)} + scores.update({"a:hog": 50.0, "a:cheap": 1.0}) + listing = engine.compute_listing( + entries, engine.ListingConfig(context_window_tokens=200_000), scores + ) + by_name = {r["qualified_name"]: r for r in listing["skills"]} + # The name floor takes 67, leaving 7933. 5 x 1536 = 7680 granted, leaving + # 253. a:hog needs 1536 and is shed; the walk does NOT stop there, and + # a:cheap needs only 200, so it is granted from the same leftovers. + self.assertEqual(by_name["a:hog"]["verdict"], "likely-starved") + self.assertEqual(by_name["a:cheap"]["verdict"], "likely-retained") + self.assertEqual(by_name["a:fill0"]["verdict"], "likely-retained") + # And the band still ranks by exposure, so the cheap survivor is band 1 + # even though it was retained. Band and verdict answer different + # questions and are allowed to disagree. + self.assertEqual(by_name["a:cheap"]["band"], 1) + class JoinerCharsTest(unittest.TestCase): """The listing inserts a literal ' - ' between description and when_to_use.