Skip to content

feat(claude-ops): inventory built-in agents and tools as native-surface lanes - #5467

Merged
kyle-sexton merged 20 commits into
mainfrom
feat/inventory-agents-tools-lanes
Sep 30, 2026
Merged

kyle-sexton merged 20 commits into
mainfrom
feat/inventory-agents-tools-lanes

Conversation

@kyle-sexton

Copy link
Copy Markdown
Contributor

No related issue: operator decision Q10 from the 2026-09-29 native-surfaces interview (follow-up to #5371).

Summary

The inventory covered commands, bundled skills and workflows, but not built-in subagent types or built-in tools. Overlaps such as discovery:explore vs the Explore agent, or planning:plan vs Plan, were invisible to native-overlap detection.

Stacked on #5371; merge that first.

Fix

  • builtin_agents lane: 11 agents.
    • Found by shape: an object with agentType, source:"built-in", and whenToUse or getSystemPrompt. Names come from literals or resolved constants.
    • Roster status: read from the roster function, as default, conditional or absent.
    • Canaries: general-purpose, Explore, Plan, statusline-setup.
  • builtin_tools lane: 80 named tools.
    • Found by shape: name plus maxResultSizeChars. The builder's minified name is never needed.
    • Names: resolved from nearest-binding constants, preferring PascalCase.
    • Deferred marker: where determinable.
    • Not counted as tools: factory-built tools are counted separately, and MCP templates are skipped.
    • Canaries: Bash, Read, Edit, Write, WebFetch.
  • Honest nulls: descriptions built by functions stay null and are counted under integrity.undetermined.
  • --docs: adds a tools cross-check against the tools reference, as its own block with its own status.
  • overlap.py detect: scores both lanes as classes builtin-agent and builtin-tool, route only. Each lane is presence-checked.
  • Native-references convention: gains the rows for both classes (3.2.0).
  • Speed: _nearest_binding runs about 35x faster, so the whole run stays at about 11.5 s.
  • The four existing lanes are byte-identical to before.
  • claude-ops 0.66.0.

Verification

  • inventory.py --self-check: all six lanes ok. It exits 3 (degraded) only on the CLI version advisory, 2.1.285 against a validated 2.1.284.

  • test_inventory.py: 128 tests OK (26 new).

  • test_overlap.py: 137 tests OK. overlap.test.sh passes, and overlap.py generate --check is in sync.

  • check-changed-skills.sh origin/main: 17 skills, 0 failed.

  • check-changelog-parity.sh: --check, --check-order, --check-bump origin/main and --check-preserved origin/main all pass.

  • Pinned ruff: check and format --check clean. typos: clean on the changed files.

  • Detect on the new lanes: 20 candidates. Top ones:

    • Plan → planning:plan (0.69)
    • Explore → discovery:explorer (0.60)
    • Explore → discovery:explore (0.57)

    They are ruled in a follow-up triage.

Related

🤖 Generated with Claude Code

kyle-sexton and others added 15 commits September 29, 2026 13:53
The brace reader treated template-literal ${...} substitutions as text, so
quotes inside a regex in a substitution desynchronized it and one brace pair
swallowed 21 MB of the bundle; 15 of 152 commands resolved. Substitutions are
now tokenized as code. Also resolves registerSlidesSkill, literal-table skill
rosters, and constant-named commands; tightens registrar and registration-token
matching instead of widening thresholds.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
detect now scores every native surface (builtin commands, bundled skills,
plugin-backed built-ins, and bundled workflows when the inventory carries
that lane) against every repo skill and agent from name and description
tokens, and emits pairs over a threshold, top-k per surface, as
origin "discovered" beside the seeded pairs. Pairs already in the store are
listed as existing with their verdict; seeds absorb their discovered twin.

Each candidate carries invocable_by from model_invocable/user_invocable
(older inventories degrade to unknown) and a recommended_integration label;
model_invocable false sets the model-invocation-disabled marker the store's
suggest-only rule reads. bundled-workflow joins the provenance classes.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Seven pairs whose overlap is conceptual rather than lexical (recap, fork,
subtask, batch, explain-usage x2, fewer-permission-prompts) score below the
discovery cut against Claude Code 2.1.284, so they join the seeded pairs.
Seeded candidates now report their lexical score even below the cut.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The report structure now shows each candidate's origin, score, invocable_by
and recommended integration, and the detection posture states how discovery
scores and where seeds still earn their place. The plugin_backed lane and
code-review alias gotchas are re-verified against Claude Code 2.1.284.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…cs cross-check

The inventory now emits argument_hint and description resolved from
getters, constants, function references, and concatenations; user_invocable
and model_invocable on every command and bundled skill, null when the
bundle decides at runtime; a bundled_workflows lane with a deep-research
canary; and a --docs mode that classifies each name against the commands
page and attaches changelog history as a labeled heuristic.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…s cross-check

SKILL.md gains the Invocable-by marker, the bundled workflows and docs
cross-check report sections, the --docs flags, and a Next pointer to the
native-overlap audit. extraction.md covers the field resolver, the
invocability rules, the workflow push-site registrar, and the docs lane.
Verification records re-checked against Claude Code 2.1.284 and the
2026-09-29 docs.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…-inventory-2.1.284

# Conflicts:
#	plugins/claude-ops/.claude-plugin/plugin.json
#	plugins/claude-ops/CHANGELOG.md
Security review of #5371: the commands-table row regex backtracks
polynomially and the link regex quadratically on a pathological line, and
the fetch read had no size cap. Rows over 8,000 characters are skipped
(longest real row: 1,638) and bodies over 16 MB degrade the block (the
changelog is 0.84 MB).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…-inventory-2.1.284

# Conflicts:
#	plugins/claude-ops/.claude-plugin/plugin.json
#	plugins/claude-ops/CHANGELOG.md
Add builtin_agents and builtin_tools lanes to the inventory extractor, each
with its own integrity status, canaries, and floor semantics. Agent types and
tool names resolve at runtime from their constants; the agent roster marks
which types a default session registers. The binding lookup now leads with
the identifier so the regex engine keeps its literal-prefix scan.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
--docs now nests a tools block that classifies the builtin_tools lane
against the tools reference table (documented, alias, undocumented,
docs_only) with its own status, so a tools-page failure never changes the
commands verdict. --tools-docs-file reads the page from a file.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…etection

detect now reads the optional builtin_agents and builtin_tools lanes as the
builtin-agent and builtin-tool classes, presence-checked like the workflow
lane. Both are route-only in the store, the native-references convention
records the rule (3.2.0), and the generated view carries the new sections.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The inventory skill and its extraction reference describe the builtin_agents
and builtin_tools lanes, their markers, report sections, canaries, and the
tools docs cross-check, with dated verification records for the upstream
claims they rely on.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…lling

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Base automatically changed from fix/claude-ops-native-inventory-2.1.284 to main September 29, 2026 21:58
kyle-sexton and others added 4 commits September 29, 2026 18:04
…-tools-lanes

# Conflicts:
#	docs/conventions/native-references/CHANGELOG.md
#	plugins/claude-ops/.claude-plugin/plugin.json
#	plugins/claude-ops/CHANGELOG.md
#	plugins/claude-ops/skills/audit-native-overlap/SKILL.md
#	plugins/claude-ops/skills/audit-native-overlap/scripts/discover.py
#	plugins/claude-ops/skills/audit-native-overlap/scripts/overlap.py
#	plugins/claude-ops/skills/audit-native-overlap/scripts/test_overlap.py
#	plugins/claude-ops/skills/inventory/SKILL.md
#	plugins/claude-ops/skills/inventory/reference/extraction.md
#	plugins/claude-ops/skills/inventory/scripts/docs_crosscheck.py
#	plugins/claude-ops/skills/inventory/scripts/inventory.py
#	plugins/claude-ops/skills/inventory/scripts/test_inventory.py
_TOOL_ROW_RE and the commands-table _ROW_RE matched every cell with lazy
`.*?` and `[^|]*?` groups between `\s*` runs, which backtrack
super-linearly on long whitespace runs: 3.1 s for a tools row with two
125-space runs, 2.1 s for a commands row with two 1000-space runs, both
far under the 8,000-character _ROW_MAX. Both regexes now match only the
first cell; the text up to the closing pipe is split with rstrip and
rpartition, which parses the same rows from the live tools and commands
pages. The _ROW_MAX comment now says it caps the text scanned, not
backtracking. Timing tests hold pathological rows under 1 s.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…-tools-lanes

# Conflicts:
#	docs/conventions/native-references/CHANGELOG.md
#	docs/native-surfaces.md
#	plugins/claude-ops/.claude-plugin/plugin.json
#	plugins/claude-ops/CHANGELOG.md
#	plugins/claude-ops/skills/audit-native-overlap/scripts/overlap.py
Adding `--tools` took the hint to 102 characters, over the shipping-tree
argument-hint budget. The Arguments line leads with the same hint.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@kyle-sexton
kyle-sexton marked this pull request as ready for review September 30, 2026 01:28
@kyle-sexton
kyle-sexton enabled auto-merge (squash) September 30, 2026 01:28
@chatgpt-codex-connector

chatgpt-codex-connector Bot commented Sep 30, 2026 •

Copy link
Copy Markdown

Codex Review Summary

This comment shows the latest Codex review activity on this pull request.

Review Status Commit Review trigger
📝 Code Review ✅ Completed 2026-09-30T01:34:29.373512Z 6ed1b3f Draft marked ready
ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review" or "@codex security review".

Codex reacts with 👀 while any review is running, comments if it has suggestions, and reacts with 👍 once all reviews finish with no findings.

@claude

claude Bot commented Sep 30, 2026 •

Copy link
Copy Markdown
Contributor

Claude finished @kyle-sexton's task in 3m 15s —— View job


Review in progress

  • Gather context (PR diff, changed files, REVIEW.md scope split)
  • Confirm a security lane exists (claude-security-review-hosted.yml) → security findings out of scope for this lane
  • Deep review of inventory.py new agent/tool extraction (background)
  • Deep review of docs_crosscheck.py, overlap.py, and new tests (background)
  • Adversarial verification of any candidate findings
  • Post findings as inline PR comments
  • Post final summary comment

@claude

claude Bot commented Sep 30, 2026 •

Copy link
Copy Markdown
Contributor

Claude finished @kyle-sexton's task in 1m 26s —— View job


Security review — PR #5467

Scope: all 14 changed files (git diff origin/main...HEAD), with focus on the new extraction logic in inventory.py, overlap.py, discover.py, and docs_crosscheck.py — the only files with executable logic.

Skip-gate check: PR is open; no prior review found on this head; the diff has code surface, so a full review was done rather than a skip.

Instruction-surface check: the diff only adds rows/sentences to docs/conventions/native-references/README.md, CHANGELOG.md, and docs/native-surfaces.md, and expands SKILL.md prose for the two affected skills. Nothing is deleted, narrowed, or softened — no guardrail, block-list, or allowlist entry is removed — so the instruction-exception-register lens has nothing to flag here.

What I looked for:

  • Injection surfaces (command, path, template) in the new static-analysis code — none: no subprocess, eval/exec, os.system, shell=True, or dynamic imports were introduced (inventory.py new section →).
  • ReDoS on untrusted input: docs_crosscheck.py fetches remote doc pages and now adds a tools-table parser. The new _row_rest helper (L94-L112) deliberately replaces the old lazy .*? cell pattern with an anchored first-cell regex plus plain string slicing for the rest of the row — this is a backtracking fix, not a regression, and it's applied consistently to both the commands-table and new tools-table parsers.
  • SSRF/URL handling: the new TOOLS_URL constant (L57) is a hardcoded first-party URL, not derived from any PR/issue/user-controlled input, and --tools-docs-file is an operator-supplied local CLI flag, same trust model as the pre-existing --docs-file/--changelog-file.
  • Secrets/token handling: none of the changed code touches credentials, auth, or tokens.
  • GitHub Actions hardening: no workflow files changed in this PR.

Result: no security issues found. This PR is a pure static-analysis extension (parsing the local Claude Code binary and two doc pages via regex) with no new attacker-reachable input, no privilege boundary, and no execution of extracted data — it's read-only reporting throughout.
· branch feat/inventory-agents-tools-lanes

…ause

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@claude

claude Bot commented Sep 30, 2026 •

Copy link
Copy Markdown
Contributor

Claude finished @kyle-sexton's task in 40s —— View job


Security review — PR #5467

Scope: full PR diff (origin/main...HEAD, 15 files), with the incremental diff since the prior review's commit (6ed1b3f) checked directly: git diff 6ed1b3f..f324466 touches only docs/catalog.md (+1/-1), a regenerated catalog description line with no code change.

Skip-gate check: the new head (f324466) adds no executable-logic change over the already-reviewed 6ed1b3f — the sole diff is prose in docs/catalog.md describing the claude-ops plugin's skill list (adds "every built-in subagent and tool" and "then re-extract the native surface and file its drift as work items" to the inventory/changelog skill descriptions). This is a pure addition, not a deletion or narrowing of any guardrail, block-list, or allowlist, so the instruction-exception-register lens has nothing to flag, and the change has no security-relevant surface.

Result: no security issues found, consistent with the prior review at 6ed1b3f (which covered all executable logic in inventory.py, overlap.py, discover.py, and docs_crosscheck.py — the only files with code in this PR — and found none). No new findings introduced by f324466075958a48c0c3ca0acdeedaa0e2a3fcc3.
· branch feat/inventory-agents-tools-lanes

@claude

claude Bot commented Sep 30, 2026 •

Copy link
Copy Markdown
Contributor

Claude finished @kyle-sexton's task in 8m 2s —— View job


Code review — PR #5467

  • Gather context (PR diff, changed files, REVIEW.md scope split)
  • Invoke /review:code-review skill and follow its procedure
  • Review new extraction logic in inventory.py
  • Review docs_crosscheck.py, overlap.py, discover.py, and new tests
  • Adversarial static trace of candidate findings
  • Post findings as inline PR comments
  • Post final summary comment

Scope: git diff origin/main...HEAD (15 files, matches the PR's declared diff). Security is already covered by a separate lane (.github/workflows/claude-security-review-hosted.yml exists and already posted "no security issues found" on this PR), so per REVIEW.md's code-review-lane scope this review omits security findings and focuses on correctness, design, error handling, and test coverage.

Findings

Posted as an inline comment on docs_crosscheck.py:122-132:

parse_tools_table has no bound on the tools-table region, unlike parse_commands_table. parse_commands_table explicitly bounds its scan to the section between ## All commands and the next \n## heading, so it can't read past the table even when every row fails to match. parse_tools_table only stops on a non-| line after rows is already non-empty — so if the real table's rows stop matching _TOOL_ROW_RE (e.g. a docs reformat drops the backtick-wrapping around tool names), rows stays empty all the way through the table, the if rows: break guard never fires, and the loop keeps scanning past the table into whatever follows in the (up to 16 MB) fetched page, picking up the first later line that happens to match | \Name` | ... |as if it were a real row. Traced by hand against the test suite's ownTOOLS_DOCS fixture (test_inventory.py:1481-1493), which deliberately places a NotAToolrow after a## Configure toolsheading specifically to prove such content is excluded in the *healthy* case — that exclusion depends entirely on the real table already having produced at least one row first. Net effect: instead of surfacing"broken"when the tools page's format drifts (the exact case this cross-check exists to catch),build_tools_crosscheckwould see a non-emptyrowsand silently report every real tool as"undocumented"` with wrong data. Details and a concrete reproduction are in the inline comment.

No other findings cleared the high-signal bar. The rest of the new extraction logic (extract_builtin_agents, extract_builtin_tools, resolve_tool_ident, _agent_roster, the _nearest_binding/build_const_index regex rewrites, and the overlap.py/discover.py lane wiring) held up under a close read — the _nearest_binding lookbehind reordering in particular is a correct rewrite of an equivalent match (verified by hand-tracing several boundary cases), not a behavior change. Test coverage for the new lanes (test_inventory.py, test_overlap.py) is thorough; I did not independently re-run the author's cited verification (pinned ruff, test_inventory.py/test_overlap.py, overlap.test.sh, changelog-parity checks) — those are author-claimed, not verified by me, since I could not execute commands in this environment (python3 invocations require approval this session had no one to grant).
· branch feat/inventory-agents-tools-lanes

@kyle-sexton
kyle-sexton merged commit a02f442 into main Sep 30, 2026
18 checks passed
@kyle-sexton
kyle-sexton deleted the feat/inventory-agents-tools-lanes branch September 30, 2026 01:45
Comment thread plugins/claude-ops/skills/inventory/scripts/docs_crosscheck.py
kyle-sexton added a commit that referenced this pull request Sep 30, 2026
…ndidates (#5504)

No related issue: follow-up to #5467, ruling on the candidates its new
built-in agent and tool lanes surfaced (operator decisions Q9 and Q10,
2026-09-29 native-surfaces interview).

## Summary

#5467 taught the inventory to extract Claude Code's built-in subagents
and tools, and taught `overlap.py detect` to score them. That surfaced
new overlap candidates with no rulings, such as `Explore` vs
`discovery:explore`, `Plan` vs `planning:plan`, and
`WebFetch`/`WebSearch` vs `firecrawl`.

## Fix

**Verdict rows** (all `complementary`, `route`, extraction-evidence
against 2.1.285):

| Native surface | Component | Split |
|---|---|---|
| `Explore` (agent) | discovery:explore | Boundary section. Built-in:
one-shot read-only locate. Ours: persisted `EXPLORE.md`. |
| `Explore` (agent) | discovery:explorer (agent) | Registry row only. |
| `Plan` (agent) | planning:plan | Boundary section. Built-in: returns
an approach and cannot write. Ours: approval-gated, persisted PLAN.md. |
| `WebFetch` (tool) | firecrawl:firecrawl | Boundary section. Built-in:
plain unprotected pages. Ours: anti-bot or JS pages, full text on disk.
|
| `WebSearch` (tool) | firecrawl:firecrawl | Boundary section. Built-in:
titles and URLs. Ours: search plus scraped content. |

**Dismissals:** 18 dismissals, each with a one-line reason:
- 17 shared-word false positives: `Write`, `Read`, `Bash`, `PowerShell`,
`Workflow`, `worker`, `Agent`, `SendUserMessage`, `memory_read`, and
further `Explore`/`Plan` pairs.
- `/output-style` vs `animation:learn-style`.

**Other changes:**
- The Boundary sections carry four-part records checked against the raw
`sub-agents.md` and `tools-reference.md` pages.
- No frontmatter description is edited; that belongs to the per-plugin
phrase sweep.
- Version bumps: discovery 0.25.11, planning 0.47.3, firecrawl 0.5.20.
The native-references Adopters table and a CHANGELOG patch entry are
included.

## Verification

- `overlap.py detect` on the 2.1.285 inventory: 0 new candidates in
every lane, 76 suppressed, 0 resurfaced, 0 orphaned.
- `overlap.py self-check`: degraded on the 2 documented advisories only,
63 rows checked. `generate --check`: in sync.
- `test_overlap.py`: 175 tests OK.
- `check-changed-skills.sh origin/main`: 0 failed. Its warnings are
outside the new sections.
- `validate-plugin-contracts.mjs`: 0 warnings. `generate-catalog.mjs
--check`: in sync.
- `check-changelog-parity.sh`: all four modes pass.
- `check-spoke-plugin-root.sh`: clean. typos: clean.

## Related

- #5467 added the lanes. #5466 added the dismissal mechanism.
- #5503 is sweep unit 1. It also touches the store and the
native-references CHANGELOG; whichever merges second takes the next
version.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

---------

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
kyle-sexton added a commit that referenced this pull request Sep 30, 2026
…lidate its inputs (#5525)

No related issue: follow-up to late review findings on the merged #5465
(Codex P1 and P2) and #5467 (Claude correctness).

## Summary

- **P1 (#5465).** `/claude-ops:changelog apply` Phase 7 treated every
`overlap.py self-check` exit 3 as passing. With no overlap store,
self-check exits 3 and declares the run report-only, but the phase still
filed work items and promoted its summary to the baseline. That could
happen in a foreign repository, or after the store went missing.
- **P2 (#5465).** `native_drift.py` accepted any syntactically valid
JSON. A wrong-shaped input, such as a `[]` summary, crashed with a
traceback and exit 1 instead of the documented exit-2 input error.
- **#5467.** `parse_tools_table` did not bound its scan. If the docs
page reformatted the tools table, the parser absorbed unrelated
backticked rows from later sections. It then reported every real tool as
`undocumented` instead of reporting itself `broken`.

## Fix

- **Report-only mode:** `native_drift.py diff` sets `"report_only":
true` when `--store` is omitted or names no file, the same test
`overlap.py` self-check uses. Items move to `unfiled`, nothing is filed,
and the baseline is kept. `context/native-drift.md` branches on that
field, not on the exit code.
- **Input shape:** `load()` checks each input kind's top-level and
nested shape (`summary`, `detect`, `store`, `inventory`) and exits 2
with `malformed <kind> input <path>: …`.
- **Tools table:** `parse_tools_table` stops at the first heading after
the table header, whether or not any row matched.
- claude-ops 0.70.1.

## Verification

- `native_drift.test.sh`: 33 tests OK. New tests cover report-only
filing nothing, a present store filing normally, and wrong-shaped inputs
for every flag; they fail on the old code.
- `test_inventory.py`: 133 tests OK. The reviewer's reproduction now
yields `{}` and a `broken` block.
- The live tools page still gives 45 documented, 35 undocumented and 1
docs_only, the same as before.
- `changelog-status.test.sh`: 70/70.
- Pinned ruff: clean. `check-changed-skills.sh origin/main`: 0 failed.
- `check-spoke-plugin-root.sh`, `validate-plugin-contracts.mjs` (0
warnings) and `generate-catalog.mjs --check`: all pass.
- `check-changelog-parity.sh`: all four modes pass. typos: clean.

## Related

- #5465 and #5467: the Codex and Claude threads there cite this PR.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

---------

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant