Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
15 commits
Select commit Hold shift + click to select a range
5b2e4ce
fix(claude-ops): repair inventory extraction for Claude Code 2.1.284
kyle-sexton Sep 29, 2026
5f15d47
feat(claude-ops): discover native overlap candidates dynamically
kyle-sexton Sep 29, 2026
1ad73a5
feat(claude-ops): seed conceptual native overlaps and score every seed
kyle-sexton Sep 29, 2026
194457d
docs(claude-ops): describe seeded and discovered overlap candidates
kyle-sexton Sep 29, 2026
4ffcf8c
feat(claude-ops): resolve built-in hints, invocability, workflows, do…
kyle-sexton Sep 29, 2026
4bfcd7c
docs(claude-ops): document inventory invocability, workflows, and doc…
kyle-sexton Sep 29, 2026
2e1e399
chore(claude-ops): release 0.64.0
kyle-sexton Sep 29, 2026
0223c23
Merge remote-tracking branch 'origin/main' into fix/claude-ops-native…
kyle-sexton Sep 29, 2026
3109884
fix(claude-ops): bound fetched docs text in the inventory cross-check
kyle-sexton Sep 29, 2026
d98c225
Merge remote-tracking branch 'origin/main' into fix/claude-ops-native…
kyle-sexton Sep 29, 2026
deb3c11
fix(claude-ops): spell a test name the way typos expects
kyle-sexton Sep 29, 2026
64a5c00
fix(claude-ops): score a plugin-backed command once and degrade on tr…
kyle-sexton Sep 29, 2026
a0dadc9
Merge remote-tracking branch 'origin/main' into fix/claude-ops-native…
kyle-sexton Sep 29, 2026
9d79ec3
fix(claude-ops): recommend route, never wrap, for a model-invocable b…
kyle-sexton Sep 29, 2026
66059c4
Merge remote-tracking branch 'origin/main' into fix/claude-ops-native…
kyle-sexton Sep 29, 2026
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
9 changes: 9 additions & 0 deletions docs/conventions/native-references/CHANGELOG.md
Original file line number Diff line number Diff line change
Expand Up @@ -6,6 +6,15 @@ major change; additive guidance is minor; clarification is a patch. The doc ship
unnumbered, which this file reads as **1.0**; the entry below is the first recorded change and lands
the changelog the README said would arrive with it.

## [3.2.0] - 2026-09-29

Minor: additive guidance.

- **`bundled-workflow` is a provenance class.** Claude Code bundles workflows (`deep-research`)
beside its skills, and the overlap store can now record a row against one. Its runtime
relationship is `route` or `suggest`, never `wrap`, because the Native step invokes through the
Skill tool and a workflow is not a skill.

## [3.1.0] - 2026-09-29

Minor: additive guidance in the Runtime relationship section.
Expand Down
1 change: 1 addition & 0 deletions docs/conventions/native-references/README.md
Original file line number Diff line number Diff line change
Expand Up @@ -236,6 +236,7 @@ at runtime. Every store row carries one of `route`, `wrap`, or `suggest`.
| `builtin-command` | `route` or `suggest`. Never `wrap`: a built-in command is not invoked by the Skill tool |
| `bundled-skill` | `route` or `wrap` |
| `bundled-skill` carrying `model-invocation-disabled` | `suggest` only. The model never lists the surface, so a route phrase is dead text. The marker is set from the registration the row's evidence names, never from the bare name |
| `bundled-workflow` | `route` or `suggest`. Never `wrap`: the Native step invokes through the Skill tool, and a workflow is not a skill |
| `plugin-backed-builtin` | `route` or `wrap` |
| `marketplace-plugin` | `route` or `wrap`. The wrap grammar for this class is seam-phrasing's, not the Native step below |
| `session-skill` | `route` only |
Expand Down
5 changes: 5 additions & 0 deletions docs/native-surfaces.md
Original file line number Diff line number Diff line change
Expand Up @@ -19,6 +19,7 @@ and when. See [`docs/conventions/native-references/`](conventions/native-referen
|---|---|---|---|---|
| Built-in CLI commands | 5 | 5 | route 1, suggest 4 | complementary 5 |
| Bundled skills | 13 | 12 | route 8, suggest 3, wrap 2 | complementary 12, defer 1 |
| Bundled workflows | 0 | 0 | none | none |
| Plugin-backed built-ins | 1 | 1 | route 1 | complementary 1 |
| Session-provided skills (observation-only) | 1 | 0 | route 1 | defer 1 |
| First-party marketplace plugins | 2 | 2 | route 2 | complementary 2 |
Expand Down Expand Up @@ -335,6 +336,10 @@ and when. See [`docs/conventions/native-references/`](conventions/native-referen
- **Baked:** description phrase yes · Boundary section yes · Native step no · suggest sentence no
- **Budget caveat:** the baked phrase may be dropped from the skill listing under budget pressure. It is the best available routing surface, not a guaranteed one

## Bundled workflows

No rows recorded in this lane.

## Plugin-backed built-ins

### `security-review` → `review:security-review`
Expand Down
2 changes: 1 addition & 1 deletion plugins/claude-ops/.claude-plugin/plugin.json
Original file line number Diff line number Diff line change
@@ -1,7 +1,7 @@
{
"$schema": "https://json.schemastore.org/claude-code-plugin-manifest.json",
"name": "claude-ops",
"version": "0.65.0",
"version": "0.66.0",
"description": "Claude Code operations toolkit. Thirteen skills: audit-skill-visibility (audit whether each installed skill is actually VISIBLE to the model, and diagnose why most of a fleet never gets used: a skill is invisible when its description is dropped by Claude Code's skill-listing context budget, which sheds descriptions lowest-score-first so an unused skill loses the keywords that would let it be matched, from skills genuinely not wanted, from skills the run cannot observe at all; computes whether the listing overflows from documented settings, and withholds every cold verdict the data cannot support rather than reporting absence of data as absence of use), inventory (read-only enumeration of the complete invocable surface: every built-in CLI command with aliases and hidden/gated status, every bundled skill, and every component of every installed plugin across all marketplaces; reads the shipped binary because upstream publishes no built-in command list, and carries an integrity verdict so a drifted build reports counts as floors rather than silently short totals), audit-install-state (read-only audit of the machine-scope ~/.claude installation directory and ~/.claude.json: full inventory split into an authored surface and rolled-up bulk trees, product-managed retention vs genuinely unmanaged state, filename-scheme resolution before any process-liveness check, and deliberate/mid-experiment detection; reports, never deletes), audit-performance (read-only slowness-diagnostic capture run at the moment the machine or a session feels slow: CLI version, retention-sweep health including the unparsable-settings pause, which warns in /status, a timed census walk of the install tree as a sweep-cost proxy, active-session and plugin-fleet counts, a process census, and the fan-out layer, which covers a load-labeled no-op spawn baseline, every hook that will fire bucketed per-tool-call versus per-turn with its invocation shape, the configured statusline, subagent concurrency and spawn-depth ceilings against documented defaults, whether running sessions predate the settings file they are judged by, and orphan attribution by parent liveness rather than age, plus on Windows a kernel-object census (Token objects against uptime, paged pool) that names a host-level leak beneath all four suspects; read against a bundled known-performance-issues reference that also records the causes tested and cleared; separates the four documented suspects of accumulated state, version regression, component bloat, and per-spawn fan-out cost, and routes remediation out; reports, never mutates, and never executes a discovered hook or statusline command), audit-native-overlap (map native Claude Code surfaces, namely built-in CLI commands, bundled skills, plugin-backed built-ins, and session-provided skills, against the current repo's plugin skills and agents, so a custom component never silently duplicates what Claude Code itself ships; bare invocation is a read-only overlap report carrying the extraction's integrity floors and a shared-listing-budget exposure section, verdicts are human-gated in a committed store rendered into a generated registry whose every row carries an observable recheck trigger, and only an explicit apply step bakes presence-gated native references into descriptions and Boundary sections), observability (read locally captured telemetry from the OTEL store, the collector, the per-session hook event log and hook-event JSONL, and ccusage, with trend reports, a per-session report of what fired, what was blocked and the event timeline, and store pruning), known-issues (search known Claude product GitHub bugs, check service health, maintain a persistent tracked-issue registry), changelog (ingest Claude Code changelog entries and integrate them into the current repo), prerequisites (read-only table of external binaries declared by enabled plugins; never installs), plugins (bring a machine's plugin fleet current on demand: marketplace refresh, effective-scope updates including in-repo project/local installs, new-plugin install per policy, scope-divergence detection and explicit convergence), morning-brief (read-only gh-based operator morning view: queue-label counts, merge-ready PRs, parked decisions with their RECOMMENDED lines, and loop-lane telemetry freshness), lanes (start/restart/stop/status loop lanes as named background Claude Code sessions seeded from canonical prompt files, with per-lane model/effort, a repo-pull + marketplace-refresh launch step, and a consume-restarts action, an OS-schedulable reader that relaunches stopped lanes whose telemetry carries a restart_request), and a re-runnable setup action that settles where the known-issues registry, the skill-usage log and the hook log root live, places the root's self-ignoring guard, and detects retired conventions. Plus an opt-in, default-off per-session hook event log (one JSON line per hook event on every event the generated registry marks observable, written to <root>/sessions/<session_id>.jsonl, with SessionEnd retention by session count or age and an optional detached pre-prune command), a family of eight advisory *-audit hooks (API errors, config changes, instruction loads, permission denials, pre-compaction, skill usage, tool failures, and unsurfaced hook failures. The last also warns the user via systemMessage, since a hook that fails to launch enforces nothing and Claude Code surfaces the failure to nobody) that emit the shared hook-telemetry envelope, and a reference sink that routes envelopes under the same root: per session when the envelope carries a session id, else into the shared hook-events.jsonl the observability skill reads.",
"author": {
"name": "Melodic Software",
Expand Down
35 changes: 35 additions & 0 deletions plugins/claude-ops/CHANGELOG.md
Original file line number Diff line number Diff line change
Expand Up @@ -3,6 +3,41 @@
All notable changes to the `claude-ops` plugin are documented here. Format follows
[Keep a Changelog](https://keepachangelog.com/en/1.1.0/); this plugin uses semantic versioning.

## [0.66.0] - 2026-09-29

### Added

- **`inventory` reports every built-in surface's arguments and invocability.** Commands, bundled
skills and bundled workflows carry `argument_hint`, `user_invocable` and `model_invocable`,
resolved from the bundle; a value the bundle computes at runtime is `null` and counted under
`integrity.undetermined`, never guessed.
- **`inventory` has a bundled-workflows lane** (`bundled_workflows`, canary `deep-research`) with
its own integrity status.
- **`inventory --docs`** fetches the live commands page and changelog and classifies each name:
documented, undocumented, alias, removed in the docs, removed but still registered, or
docs-only, with kind and alias disagreements. A fetch failure degrades that block only.
`--docs-file` and `--changelog-file` run it offline.
- **`audit-native-overlap detect` discovers candidates** by scoring every native surface against
every repo skill and agent (`--threshold`, `--top-k`), beside the seeded pairs. Each candidate
carries `invocable_by` and a `recommended_integration` label: a user-only surface is
recommended as `suggest`. A label is never a verdict.

### Fixed

- **`inventory` extraction on Claude Code 2.1.284.** Template-literal substitutions are now read as
code, so a quote inside a regex in `${...}` no longer desynchronizes the brace reader (15 of 152
commands resolved before). `registerSlidesSkill`, literal-table skill rosters and
constant-named commands resolve. Validated against 2.1.284.
- **`inventory --docs` bounds untrusted text.** A fetched body over 16 MB degrades the docs block
instead of loading, and a table row over 8,000 characters is skipped, so a malformed page
cannot stall the parser on regex backtracking. A body truncated after its headers
(`http.client.HTTPException`) degrades the block instead of raising.
- **`audit-native-overlap detect` scores a plugin-backed command once**, under its plugin-backed
class and with the description the extractor enriched it with, instead of adding a bare second
surface.
- **A model-invocable bundled workflow is recommended `route`, never `wrap`**, matching the store
rule that rejects `wrap` on a bundled-workflow row.

## [0.65.0] - 2026-09-29

### Added
Expand Down
57 changes: 44 additions & 13 deletions plugins/claude-ops/skills/audit-native-overlap/SKILL.md
Original file line number Diff line number Diff line change
Expand Up @@ -84,7 +84,10 @@ python3 "${CLAUDE_PLUGIN_ROOT}/skills/audit-native-overlap/scripts/overlap.py" s
```

`--repo`, `--store`, `--view`, and `--pairs` are flags with repo-relative defaults, so a consumer
repository with a different layout points them wherever its files live.
repository with a different layout points them wherever its files live. `detect` also takes
`--threshold` (lowest discovery score kept, default `0.30`) and `--top-k` (most components kept per
native surface, default `3`; `0` turns discovery off). Lower the threshold to `0.25` for a
recall-first sweep; below that, most added pairs share one incidental word.

Every subcommand exits `0` ok, `1` broken, `3` degraded, the sibling extractor's contract, not the
shell gates' `0/1/2`, so one lane can carry both; `2` stays argparse's usage error. A degraded exit
Expand All @@ -94,10 +97,35 @@ for the repo's test discovery.

## Detection posture. Floor-honest

Under-recall stated honestly beats confident completeness. Three rules:
Under-recall stated honestly beats confident completeness. Candidates come from two origins:

- **Seeded**: every pair in `reference/canonical-pairs.json`, emitted whether or not the extraction
shows the native side.
- **Discovered**: every native surface in the extraction, internal ones excepted, scored against
every skill and agent in the repo from name, alias, and description tokens
(`scripts/discover.py` states the formula). A pair at or over `--threshold`, among the `--top-k`
best for its surface, is emitted with its score and the shared tokens as evidence. A pair the
store already records is listed under `discovery.existing` with its verdict instead; a pair that
is also seeded stays one seeded candidate carrying the score.

Discovery is lexical. A one-line native description rarely shares words with the component it
duplicates in concept (`recap` against `session-flow:orient`), so such a pair belongs in the seeds
once a human confirms it; a pair discovery already finds needs no seed.

Every candidate carries `native.invocable_by` (`model+user`, `user-only`, `model-only`, or
`unknown`), read from the registration's `model_invocable` and `user_invocable` fields, or from
`disable_model_invocation` in an older extraction; a field the extraction lacks makes it `unknown`.
`model_invocable: false` also sets the `model-invocation-disabled` marker the store's suggest-only
rule reads. From it comes `recommended_integration`: `suggest` for a user-only surface, `route` or
`route-or-wrap` for a model-invocable one (`route` for a built-in command or a bundled workflow,
which never take `wrap`), and nothing when unknown. It is a label for the human writing the row, never a store value.

Three rules:

- **Carry the integrity floor through, per lane.** The inventory reports integrity per lane
(`builtin_commands`, `bundled_skills`, `plugin_backed`). A `degraded` lane makes every count from
(`builtin_commands`, `bundled_skills`, `plugin_backed`, and `bundled_workflows` when the
extraction has that lane; an extraction without it is not an error, the lane is simply not
reported). A `degraded` lane makes every count from
that lane a floor, and the report says so in the same sentence as the number. A `broken` lane's
counts are omitted, the report names the lane and its cause, and every candidate whose lane is
broken is marked `re_derivable: false` (its presence or absence in that lane proves nothing
Expand All @@ -118,9 +146,11 @@ Inventory status per lane (ok | degraded | broken), cli_version vs validated_aga
that means for every count below; a broken lane is named with its cause.

## Overlap candidates
One row per (native surface, our component): native name + provenance class + hidden/gated
markers, our component, the evidence, and the store's current verdict — or NEW where the
store has no row yet.
One row per (native surface, our component): origin (seeded | discovered), native name +
provenance class + hidden/gated markers, invocable_by, our component, score and shared tokens,
recommended integration (a label, not a verdict), the evidence, and the store's current verdict,
or NEW where the store has no row yet. Discovered pairs the store already records follow as one
line each with their verdict.

## Registry state
Rows whose recheck trigger has fired, rows missing a baked line, rows baked but unverified.
Expand All @@ -133,8 +163,8 @@ to name-only degradation.
Which substrate produced which section, and anything the run could not resolve.
```

Provenance classes are never merged into one list. A bundled skill, a built-in command, a
plugin-backed built-in, and a session-provided skill have different disable switches and different
Provenance classes are never merged into one list. A bundled skill, a bundled workflow, a built-in
command, a plugin-backed built-in, and a session-provided skill have different disable switches and different
rosters per host; a merged list cannot be acted on.

## Budget exposure, a presence-gated seam
Expand Down Expand Up @@ -308,14 +338,15 @@ Two upstream facts this skill depends on, each with the trigger that obliges re-
- **A plugin skill never shadows a native one.** Ours are namespaced, so both resolve and the model
chooses. That is why the routing lives in descriptions rather than in a name.
- **`plugin_backed` is its own lane.** `security-review` is reported there, not under
`builtin_commands`. Read the wrong key and the row looks absent. Verified 2026-09-06 against
Claude Code 2.1.263, by running `inventory.py --binary-only` on this machine and reading the
`plugin_backed` key, which holds `security-review` and nothing else. Recheck when the extractor's
provenance lanes change or a release note moves a bundled surface between them.
`builtin_commands`. Read the wrong key and the row looks absent. Verified 2026-09-29 against
Claude Code 2.1.284, by reading the `plugin_backed` key of an `inventory.py --binary-only`
extraction on this machine, which holds `security-review` and nothing else. Recheck when the
extractor's provenance lanes change or a release note moves a bundled surface between them.
- **A bundled skill can carry aliases.** `code-review` answers to `review`; treating an alias as a
separate surface produces a duplicate row for one capability. Basis:
<https://code.claude.com/docs/en/commands> gives `/code-review` the line "Alias: `/review`".
Verified 2026-09-06 against Claude Code 2.1.263 and that page as fetched that day. Recheck when
Verified 2026-09-29 against Claude Code 2.1.284 (the extraction lists `review` as the alias) and
that page as fetched that day. Recheck when
the commands page drops the alias line or a release note renames a bundled skill.
- **Absent from the binary is not absent from the product.** Session-provided skills exist only in
a live roster. "Not in the extraction" is a statement about the extraction.
Expand Down
Loading
Loading