feat(claude-ops): inventory built-in agents and tools as native-surface lanes - #5467
Conversation
The brace reader treated template-literal ${...} substitutions as text, so
quotes inside a regex in a substitution desynchronized it and one brace pair
swallowed 21 MB of the bundle; 15 of 152 commands resolved. Substitutions are
now tokenized as code. Also resolves registerSlidesSkill, literal-table skill
rosters, and constant-named commands; tightens registrar and registration-token
matching instead of widening thresholds.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
detect now scores every native surface (builtin commands, bundled skills, plugin-backed built-ins, and bundled workflows when the inventory carries that lane) against every repo skill and agent from name and description tokens, and emits pairs over a threshold, top-k per surface, as origin "discovered" beside the seeded pairs. Pairs already in the store are listed as existing with their verdict; seeds absorb their discovered twin. Each candidate carries invocable_by from model_invocable/user_invocable (older inventories degrade to unknown) and a recommended_integration label; model_invocable false sets the model-invocation-disabled marker the store's suggest-only rule reads. bundled-workflow joins the provenance classes. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Seven pairs whose overlap is conceptual rather than lexical (recap, fork, subtask, batch, explain-usage x2, fewer-permission-prompts) score below the discovery cut against Claude Code 2.1.284, so they join the seeded pairs. Seeded candidates now report their lexical score even below the cut. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The report structure now shows each candidate's origin, score, invocable_by and recommended integration, and the detection posture states how discovery scores and where seeds still earn their place. The plugin_backed lane and code-review alias gotchas are re-verified against Claude Code 2.1.284. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…cs cross-check The inventory now emits argument_hint and description resolved from getters, constants, function references, and concatenations; user_invocable and model_invocable on every command and bundled skill, null when the bundle decides at runtime; a bundled_workflows lane with a deep-research canary; and a --docs mode that classifies each name against the commands page and attaches changelog history as a labeled heuristic. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…s cross-check SKILL.md gains the Invocable-by marker, the bundled workflows and docs cross-check report sections, the --docs flags, and a Next pointer to the native-overlap audit. extraction.md covers the field resolver, the invocability rules, the workflow push-site registrar, and the docs lane. Verification records re-checked against Claude Code 2.1.284 and the 2026-09-29 docs. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…-inventory-2.1.284 # Conflicts: # plugins/claude-ops/.claude-plugin/plugin.json # plugins/claude-ops/CHANGELOG.md
Security review of #5371: the commands-table row regex backtracks polynomially and the link regex quadratically on a pathological line, and the fetch read had no size cap. Rows over 8,000 characters are skipped (longest real row: 1,638) and bodies over 16 MB degrade the block (the changelog is 0.84 MB). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…-inventory-2.1.284 # Conflicts: # plugins/claude-ops/.claude-plugin/plugin.json # plugins/claude-ops/CHANGELOG.md
Add builtin_agents and builtin_tools lanes to the inventory extractor, each with its own integrity status, canaries, and floor semantics. Agent types and tool names resolve at runtime from their constants; the agent roster marks which types a default session registers. The binding lookup now leads with the identifier so the regex engine keeps its literal-prefix scan. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
--docs now nests a tools block that classifies the builtin_tools lane against the tools reference table (documented, alias, undocumented, docs_only) with its own status, so a tools-page failure never changes the commands verdict. --tools-docs-file reads the page from a file. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…etection detect now reads the optional builtin_agents and builtin_tools lanes as the builtin-agent and builtin-tool classes, presence-checked like the workflow lane. Both are route-only in the store, the native-references convention records the rule (3.2.0), and the generated view carries the new sections. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The inventory skill and its extraction reference describe the builtin_agents and builtin_tools lanes, their markers, report sections, canaries, and the tools docs cross-check, with dated verification records for the upstream claims they rely on. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…lling Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…-tools-lanes # Conflicts: # docs/conventions/native-references/CHANGELOG.md # plugins/claude-ops/.claude-plugin/plugin.json # plugins/claude-ops/CHANGELOG.md # plugins/claude-ops/skills/audit-native-overlap/SKILL.md # plugins/claude-ops/skills/audit-native-overlap/scripts/discover.py # plugins/claude-ops/skills/audit-native-overlap/scripts/overlap.py # plugins/claude-ops/skills/audit-native-overlap/scripts/test_overlap.py # plugins/claude-ops/skills/inventory/SKILL.md # plugins/claude-ops/skills/inventory/reference/extraction.md # plugins/claude-ops/skills/inventory/scripts/docs_crosscheck.py # plugins/claude-ops/skills/inventory/scripts/inventory.py # plugins/claude-ops/skills/inventory/scripts/test_inventory.py
_TOOL_ROW_RE and the commands-table _ROW_RE matched every cell with lazy `.*?` and `[^|]*?` groups between `\s*` runs, which backtrack super-linearly on long whitespace runs: 3.1 s for a tools row with two 125-space runs, 2.1 s for a commands row with two 1000-space runs, both far under the 8,000-character _ROW_MAX. Both regexes now match only the first cell; the text up to the closing pipe is split with rstrip and rpartition, which parses the same rows from the live tools and commands pages. The _ROW_MAX comment now says it caps the text scanned, not backtracking. Timing tests hold pathological rows under 1 s. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…-tools-lanes # Conflicts: # docs/conventions/native-references/CHANGELOG.md # docs/native-surfaces.md # plugins/claude-ops/.claude-plugin/plugin.json # plugins/claude-ops/CHANGELOG.md # plugins/claude-ops/skills/audit-native-overlap/scripts/overlap.py
Adding `--tools` took the hint to 102 characters, over the shipping-tree argument-hint budget. The Arguments line leads with the same hint. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Codex Review SummaryThis comment shows the latest Codex review activity on this pull request.
ℹ️ About Codex in GitHubYour team has set up Codex to review pull requests in this repo. Reviews are triggered when you
Codex reacts with 👀 while any review is running, comments if it has suggestions, and reacts with 👍 once all reviews finish with no findings. |
|
Claude finished @kyle-sexton's task in 3m 15s —— View job Review in progress
|
|
Claude finished @kyle-sexton's task in 1m 26s —— View job Security review — PR #5467
Scope: all 14 changed files ( Skip-gate check: PR is open; no prior review found on this head; the diff has code surface, so a full review was done rather than a skip. Instruction-surface check: the diff only adds rows/sentences to What I looked for:
Result: no security issues found. This PR is a pure static-analysis extension (parsing the local Claude Code binary and two doc pages via regex) with no new attacker-reachable input, no privilege boundary, and no execution of extracted data — it's read-only reporting throughout. |
…ause Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
|
Claude finished @kyle-sexton's task in 40s —— View job Security review — PR #5467
Scope: full PR diff ( Skip-gate check: the new head ( Result: no security issues found, consistent with the prior review at |
|
Claude finished @kyle-sexton's task in 8m 2s —— View job Code review — PR #5467
Scope: FindingsPosted as an inline comment on
No other findings cleared the high-signal bar. The rest of the new extraction logic ( |
…ndidates (#5504) No related issue: follow-up to #5467, ruling on the candidates its new built-in agent and tool lanes surfaced (operator decisions Q9 and Q10, 2026-09-29 native-surfaces interview). ## Summary #5467 taught the inventory to extract Claude Code's built-in subagents and tools, and taught `overlap.py detect` to score them. That surfaced new overlap candidates with no rulings, such as `Explore` vs `discovery:explore`, `Plan` vs `planning:plan`, and `WebFetch`/`WebSearch` vs `firecrawl`. ## Fix **Verdict rows** (all `complementary`, `route`, extraction-evidence against 2.1.285): | Native surface | Component | Split | |---|---|---| | `Explore` (agent) | discovery:explore | Boundary section. Built-in: one-shot read-only locate. Ours: persisted `EXPLORE.md`. | | `Explore` (agent) | discovery:explorer (agent) | Registry row only. | | `Plan` (agent) | planning:plan | Boundary section. Built-in: returns an approach and cannot write. Ours: approval-gated, persisted PLAN.md. | | `WebFetch` (tool) | firecrawl:firecrawl | Boundary section. Built-in: plain unprotected pages. Ours: anti-bot or JS pages, full text on disk. | | `WebSearch` (tool) | firecrawl:firecrawl | Boundary section. Built-in: titles and URLs. Ours: search plus scraped content. | **Dismissals:** 18 dismissals, each with a one-line reason: - 17 shared-word false positives: `Write`, `Read`, `Bash`, `PowerShell`, `Workflow`, `worker`, `Agent`, `SendUserMessage`, `memory_read`, and further `Explore`/`Plan` pairs. - `/output-style` vs `animation:learn-style`. **Other changes:** - The Boundary sections carry four-part records checked against the raw `sub-agents.md` and `tools-reference.md` pages. - No frontmatter description is edited; that belongs to the per-plugin phrase sweep. - Version bumps: discovery 0.25.11, planning 0.47.3, firecrawl 0.5.20. The native-references Adopters table and a CHANGELOG patch entry are included. ## Verification - `overlap.py detect` on the 2.1.285 inventory: 0 new candidates in every lane, 76 suppressed, 0 resurfaced, 0 orphaned. - `overlap.py self-check`: degraded on the 2 documented advisories only, 63 rows checked. `generate --check`: in sync. - `test_overlap.py`: 175 tests OK. - `check-changed-skills.sh origin/main`: 0 failed. Its warnings are outside the new sections. - `validate-plugin-contracts.mjs`: 0 warnings. `generate-catalog.mjs --check`: in sync. - `check-changelog-parity.sh`: all four modes pass. - `check-spoke-plugin-root.sh`: clean. typos: clean. ## Related - #5467 added the lanes. #5466 added the dismissal mechanism. - #5503 is sweep unit 1. It also touches the store and the native-references CHANGELOG; whichever merges second takes the next version. 🤖 Generated with [Claude Code](https://claude.com/claude-code) --------- Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
…lidate its inputs (#5525) No related issue: follow-up to late review findings on the merged #5465 (Codex P1 and P2) and #5467 (Claude correctness). ## Summary - **P1 (#5465).** `/claude-ops:changelog apply` Phase 7 treated every `overlap.py self-check` exit 3 as passing. With no overlap store, self-check exits 3 and declares the run report-only, but the phase still filed work items and promoted its summary to the baseline. That could happen in a foreign repository, or after the store went missing. - **P2 (#5465).** `native_drift.py` accepted any syntactically valid JSON. A wrong-shaped input, such as a `[]` summary, crashed with a traceback and exit 1 instead of the documented exit-2 input error. - **#5467.** `parse_tools_table` did not bound its scan. If the docs page reformatted the tools table, the parser absorbed unrelated backticked rows from later sections. It then reported every real tool as `undocumented` instead of reporting itself `broken`. ## Fix - **Report-only mode:** `native_drift.py diff` sets `"report_only": true` when `--store` is omitted or names no file, the same test `overlap.py` self-check uses. Items move to `unfiled`, nothing is filed, and the baseline is kept. `context/native-drift.md` branches on that field, not on the exit code. - **Input shape:** `load()` checks each input kind's top-level and nested shape (`summary`, `detect`, `store`, `inventory`) and exits 2 with `malformed <kind> input <path>: …`. - **Tools table:** `parse_tools_table` stops at the first heading after the table header, whether or not any row matched. - claude-ops 0.70.1. ## Verification - `native_drift.test.sh`: 33 tests OK. New tests cover report-only filing nothing, a present store filing normally, and wrong-shaped inputs for every flag; they fail on the old code. - `test_inventory.py`: 133 tests OK. The reviewer's reproduction now yields `{}` and a `broken` block. - The live tools page still gives 45 documented, 35 undocumented and 1 docs_only, the same as before. - `changelog-status.test.sh`: 70/70. - Pinned ruff: clean. `check-changed-skills.sh origin/main`: 0 failed. - `check-spoke-plugin-root.sh`, `validate-plugin-contracts.mjs` (0 warnings) and `generate-catalog.mjs --check`: all pass. - `check-changelog-parity.sh`: all four modes pass. typos: clean. ## Related - #5465 and #5467: the Codex and Claude threads there cite this PR. 🤖 Generated with [Claude Code](https://claude.com/claude-code) --------- Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>

No related issue: operator decision Q10 from the 2026-09-29 native-surfaces interview (follow-up to #5371).
Summary
The inventory covered commands, bundled skills and workflows, but not built-in subagent types or built-in tools. Overlaps such as
discovery:explorevs theExploreagent, orplanning:planvsPlan, were invisible to native-overlap detection.Stacked on #5371; merge that first.
Fix
builtin_agentslane: 11 agents.agentType,source:"built-in", andwhenToUseorgetSystemPrompt. Names come from literals or resolved constants.builtin_toolslane: 80 named tools.nameplusmaxResultSizeChars. The builder's minified name is never needed.integrity.undetermined.--docs: adds a tools cross-check against the tools reference, as its own block with its own status.overlap.py detect: scores both lanes as classesbuiltin-agentandbuiltin-tool,routeonly. Each lane is presence-checked._nearest_bindingruns about 35x faster, so the whole run stays at about 11.5 s.claude-ops0.66.0.Verification
inventory.py --self-check: all six lanesok. It exits 3 (degraded) only on the CLI version advisory, 2.1.285 against a validated 2.1.284.test_inventory.py: 128 tests OK (26 new).test_overlap.py: 137 tests OK.overlap.test.shpasses, andoverlap.py generate --checkis in sync.check-changed-skills.sh origin/main: 17 skills, 0 failed.check-changelog-parity.sh:--check,--check-order,--check-bump origin/mainand--check-preserved origin/mainall pass.Pinned ruff:
checkandformat --checkclean. typos: clean on the changed files.Detect on the new lanes: 20 candidates. Top ones:
They are ruled in a follow-up triage.
Related
🤖 Generated with Claude Code