Skip to content

fix(claude-ops): repair inventory for 2.1.284 and discover native overlaps dynamically - #5371

Merged
kyle-sexton merged 15 commits into
mainfrom
fix/claude-ops-native-inventory-2.1.284
Sep 29, 2026
Merged

kyle-sexton merged 15 commits into
mainfrom
fix/claude-ops-native-inventory-2.1.284

Conversation

@kyle-sexton

Copy link
Copy Markdown
Contributor

No related issue: the inventory extractor broke on Claude Code 2.1.284, found while building a full built-in command inventory.

Summary

/claude-ops:inventory resolved 15 of 152 command registrations on Claude Code 2.1.284 and missed every canary (help, clear, config, resume, status). Its self-check reported the lane as broken rather than printing a short list. audit-native-overlap reads that inventory and could only report hand-seeded pairs, so the new 2.1.284 surfaces (bundled commit and pr, /autofix-pr, /subtask, the deep-research workflow, and others) never showed up as overlaps.

Fix

  • Extractor: template-literal ${...} substitutions are now tokenized as code. Before, a quote inside a regex inside a substitution made one brace pair span 21 MB. The extractor now also resolves registerSlidesSkill, literal-table skill rosters and constant-named commands. Validated against 2.1.284.
  • Inventory: every command, bundled skill and bundled workflow now carries argument_hint, user_invocable and model_invocable, read from the bundle. A value the bundle computes at runtime is left null and counted, never guessed. New bundled_workflows lane.
  • inventory --docs: fetches the live commands page and changelog and classifies each name: documented, undocumented, alias, removed, removed but still registered, or docs-only. Kind and alias disagreements are reported too. The binary stays the source of truth for what exists.
  • audit-native-overlap detect: scores every native surface against every repo skill and agent (--threshold 0.30, --top-k 3) alongside the seeded pairs. Each candidate carries invocable_by and a recommended_integration label, so a user-only surface is recommended as suggest. A label is never a verdict. Seven new seeds cover overlaps whose wording doesn't match.
  • claude-ops 0.64.0; native-references convention gains a bundled-workflow row (3.1.0).

Verification

  • inventory.py --self-check: OK: cli 2.1.284, validated against 2.1.284, lanes builtin_commands, bundled_skills, plugin_backed and bundled_workflows all ok.
  • test_inventory.py: 101 tests OK. test_overlap.py: 134 tests OK. overlap.test.sh passes.
  • Pinned ruff clean. check-changed-skills.sh main: 2 skills, 0 failed. check-changelog-parity.sh: --check, --check-bump main, --check-order and --check-preserved main all exit 0.
  • On 2.1.284:
    • inventory: 111 commands, 39 bundled skills, 1 plugin-backed command, 1 bundled workflow
    • docs cross-check: 108 documented, 39 undocumented, 20 aliases, 3 removed
    • detect: 106 candidates (23 seeded, 83 discovered)
  • Discovery precision is about 41% on each surface's top pick. The false positives mostly share one generic word; a person rules on every candidate.

Related

  • A follow-up PR will record verdicts for the new pairs and add Boundary sections. Description phrases will follow one plugin per PR under the sweep contract.

🤖 Generated with Claude Code

kyle-sexton and others added 7 commits September 29, 2026 13:53
The brace reader treated template-literal ${...} substitutions as text, so
quotes inside a regex in a substitution desynchronized it and one brace pair
swallowed 21 MB of the bundle; 15 of 152 commands resolved. Substitutions are
now tokenized as code. Also resolves registerSlidesSkill, literal-table skill
rosters, and constant-named commands; tightens registrar and registration-token
matching instead of widening thresholds.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
detect now scores every native surface (builtin commands, bundled skills,
plugin-backed built-ins, and bundled workflows when the inventory carries
that lane) against every repo skill and agent from name and description
tokens, and emits pairs over a threshold, top-k per surface, as
origin "discovered" beside the seeded pairs. Pairs already in the store are
listed as existing with their verdict; seeds absorb their discovered twin.

Each candidate carries invocable_by from model_invocable/user_invocable
(older inventories degrade to unknown) and a recommended_integration label;
model_invocable false sets the model-invocation-disabled marker the store's
suggest-only rule reads. bundled-workflow joins the provenance classes.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Seven pairs whose overlap is conceptual rather than lexical (recap, fork,
subtask, batch, explain-usage x2, fewer-permission-prompts) score below the
discovery cut against Claude Code 2.1.284, so they join the seeded pairs.
Seeded candidates now report their lexical score even below the cut.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The report structure now shows each candidate's origin, score, invocable_by
and recommended integration, and the detection posture states how discovery
scores and where seeds still earn their place. The plugin_backed lane and
code-review alias gotchas are re-verified against Claude Code 2.1.284.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…cs cross-check

The inventory now emits argument_hint and description resolved from
getters, constants, function references, and concatenations; user_invocable
and model_invocable on every command and bundled skill, null when the
bundle decides at runtime; a bundled_workflows lane with a deep-research
canary; and a --docs mode that classifies each name against the commands
page and attaches changelog history as a labeled heuristic.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…s cross-check

SKILL.md gains the Invocable-by marker, the bundled workflows and docs
cross-check report sections, the --docs flags, and a Next pointer to the
native-overlap audit. extraction.md covers the field resolver, the
invocability rules, the workflow push-site registrar, and the docs lane.
Verification records re-checked against Claude Code 2.1.284 and the
2026-09-29 docs.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
kyle-sexton and others added 3 commits September 29, 2026 16:34
…-inventory-2.1.284

# Conflicts:
#	plugins/claude-ops/.claude-plugin/plugin.json
#	plugins/claude-ops/CHANGELOG.md
Security review of #5371: the commands-table row regex backtracks
polynomially and the link regex quadratically on a pathological line, and
the fetch read had no size cap. Rows over 8,000 characters are skipped
(longest real row: 1,638) and bodies over 16 MB degrade the block (the
changelog is 0.84 MB).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…-inventory-2.1.284

# Conflicts:
#	plugins/claude-ops/.claude-plugin/plugin.json
#	plugins/claude-ops/CHANGELOG.md
@kyle-sexton
kyle-sexton marked this pull request as ready for review September 29, 2026 20:43
@chatgpt-codex-connector

chatgpt-codex-connector Bot commented Sep 29, 2026 •

Copy link
Copy Markdown

Codex Review Summary

This comment shows the latest Codex review activity on this pull request.

Review Status Commit Review trigger
📝 Code Review ✅ Completed 2026-09-29T20:47:35.739420Z d98c225 Draft marked ready
ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review" or "@codex security review".

Codex reacts with 👀 while any review is running, comments if it has suggestions, and reacts with 👍 once all reviews finish with no findings.

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: d98c2251a8

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

Comment thread plugins/claude-ops/skills/audit-native-overlap/scripts/overlap.py Outdated
Comment thread plugins/claude-ops/skills/inventory/scripts/docs_crosscheck.py
@claude

claude Bot commented Sep 29, 2026 •

Copy link
Copy Markdown
Contributor

Claude finished @kyle-sexton's task in 1m 53s —— View job


Claude is reviewing this PR

  • Gather context (PR body, changed files, repo instructions)
  • Invoke /review:code-review and follow its procedure
  • Get canonical diff (gh pr diff 5371, 15 files match the PR's changed-file list)
  • Review diff across extractor, discovery/overlap scoring, docs cross-check, and docs/metadata consistency (4 parallel passes running)
  • Verify candidate findings adversarially
  • Post inline findings
  • Post summary comment

@claude

claude Bot commented Sep 29, 2026 •

Copy link
Copy Markdown
Contributor

Claude finished @kyle-sexton's task in 2m 1s —— View job


CI security review (/review:security-review)

  • Gather context (PR diff, changed files)
  • Check skip gate (open PR, security-relevant surface: yes — new network fetch + text extraction logic)
  • Review extractor/discovery scripts for injection, SSRF, unsafe parsing
  • Review docs cross-check network fetch code
  • Check for instruction-surface deletions/weakenings
  • Post findings

No security issues found.

Scope: the diff's only genuinely untrusted-input surface is docs_crosscheck.py's two fetches (code.claude.com/docs/en/commands.md, raw.githubusercontent.com/anthropics/claude-code/main/CHANGELOG.md). Both URLs are hardcoded constants (no SSRF), and the PR itself adds the hardening this lane would otherwise flag: a 16 MB response cap and an 8,000-char per-row cap before the extraction regexes run, specifically to bound catastrophic-backtracking risk on untrusted fetched text (docs_crosscheck.py L63-L79).

Everything else touched is local, trusted input: the installed claude binary (read-only, never executed — inventory.py), the repo's own JSON store/pairs files and SKILL.md tree (overlap.py, discover.py), and a fixed-argv subprocess.run(["claude", "--version"], ...) with no shell (overlap.py L472-L479). No eval/exec, no shell=True, no path built from untrusted input.

Checked the diff's deletions against the instruction-exception-register lens (removed/softened guardrails in CLAUDE.md/AGENTS.md/rules/skill bodies): none. All - lines in SKILL.md/extraction.md/rules files are prose rewrites or superseded-by-new-content replacements (version bumps, restructured sections), not the removal of a standing check or control.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@claude

claude Bot commented Sep 29, 2026 •

Copy link
Copy Markdown
Contributor

Claude finished @kyle-sexton's task in 3m 3s —— View job


Claude is reviewing this PR

  • Gather context (PR body, changed files, repo instructions)
  • Invoke /review:code-review and follow its procedure
  • Get canonical diff (gh pr diff 5371, 15 files match the PR's changed-file list)
  • Review diff across extractor, discovery/overlap scoring, docs cross-check, and docs/metadata consistency (parallel passes running)
  • Verify candidate findings adversarially
  • Post inline findings
  • Post summary comment

@claude

claude Bot commented Sep 29, 2026 •

Copy link
Copy Markdown
Contributor

Claude finished @kyle-sexton's task in 27s —— View job


CI security review (/review:security-review)

  • Gather context (PR diff, changed files, repo instructions)
  • Invoke /review:security-review skill and follow its procedure exactly
  • Check skip gate
  • Post findings

Skip gate: criterion 3 (this head already has a successful security review that still applies).

This lane already reviewed the PR at commit d98c225 (see the comment above,
job run) and found no security issues, covering all 15 changed files including the fetch-hardening in docs_crosscheck.py and the subprocess call in overlap.py.

The only commit since then, deb3c11, changes exactly one line:

- def test_an_inventory_without_lanes_keeps_the_old_behaviour(
+ def test_an_inventory_without_lanes_keeps_the_old_behavior(

A unit-test method rename (US spelling for a spellchecker tool), no functional or security-relevant surface. Confirmed via git diff d98c225..HEAD — one file, one line, no code paths touched.

No security issues found. The prior review at d98c225 still applies to HEAD (deb3c11); nothing new to assess.

kyle-sexton and others added 3 commits September 29, 2026 17:25
…uncated fetches

Review fixes for #5371 (Codex P2 findings):
- detect built a second, description-free surface for a plugin-backed name
  the extractor had already enriched in the commands or skills lane, and
  the native index replaced the enriched entry with the bare one. Both now
  reclassify the enriched registration instead.
- fetch_text did not catch http.client.HTTPException, so an IncompleteRead
  mid-body crashed `--docs` instead of degrading its block.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…-inventory-2.1.284

# Conflicts:
#	docs/conventions/native-references/CHANGELOG.md
#	plugins/claude-ops/.claude-plugin/plugin.json
#	plugins/claude-ops/CHANGELOG.md
…undled workflow

The store validator rejects `wrap` on a bundled-workflow row, but
recommended_integration only special-cased built-in commands, so detect
labelled a model-invocable workflow `route-or-wrap` and led the row author
to a value self-check refuses. Found by a local /code-review standing in
for the Claude review lane on #5371.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@claude

claude Bot commented Sep 29, 2026 •

Copy link
Copy Markdown
Contributor

Claude encountered an error after 0s —— View job


I'll analyze this and get back to you.

@claude

claude Bot commented Sep 29, 2026 •

Copy link
Copy Markdown
Contributor

Claude encountered an error after 0s —— View job


I'll analyze this and get back to you.

@kyle-sexton

Copy link
Copy Markdown
Contributor Author

Review lanes on 9d79ec3ef: absent (usage limit, 429), substituted locally.

Lane CI result Substitute
review / claude-review-status failed with failure-class=rate-limit, so no review ran. The earlier runs on d98c225 and 64a5c00 stopped before posting findings. Local /code-review medium 5371 over the full PR diff. It found 1 issue: recommended_integration offered wrap for a bundled workflow, which the store rejects. Fixed in 9d79ec3ef with a test.
security-review / claude-security-review-status failed with failure-class=rate-limit The lane's own review at d98c225 found no issues. A local security review of the same diff found 2 P4 regex-backtracking issues and an uncapped fetch, all fixed in 310988495. Every commit since touches no trust boundary: a test rename, catching http.client.HTTPException, plugin-backed reclassification, and a recommendation label.

Codex: 2 P2 findings, both fixed in 64a5c0042; the threads are replied to and resolved.

…-inventory-2.1.284

# Conflicts:
#	plugins/claude-ops/CHANGELOG.md
@kyle-sexton
kyle-sexton enabled auto-merge (squash) September 29, 2026 21:53
@claude

claude Bot commented Sep 29, 2026 •

Copy link
Copy Markdown
Contributor

Claude encountered an error after 0s —— View job


I'll analyze this and get back to you.

@claude

claude Bot commented Sep 29, 2026 •

Copy link
Copy Markdown
Contributor

Claude encountered an error after 0s —— View job


I'll analyze this and get back to you.

@kyle-sexton
kyle-sexton merged commit f466855 into main Sep 29, 2026
17 of 19 checks passed
@kyle-sexton
kyle-sexton deleted the fix/claude-ops-native-inventory-2.1.284 branch September 29, 2026 21:58
kyle-sexton added a commit that referenced this pull request Sep 29, 2026
…ections (#5387)

No related issue: follow-up to #5371, recording verdicts for native
Claude Code surfaces the 2.1.284 overlap detector found.

## Summary

Claude Code 2.1.284 ships native surfaces that overlap this
marketplace's skills: bundled `commit` and `pr`, `/autofix-pr`,
`/subtask`, `/fork`, `/background`, `/recap`, the `deep-research`
workflow, `/doctor prompt-audit`, `/verify`, `/batch`, and others.
Nothing told the model when to prefer the native surface, or that a
user-only command exists to offer the person.

Stacked on #5371; merge that first.

## Fix

- **`docs/native-surfaces/records.json`:** 25 new rows (22
complementary, 3 defer), and every existing extraction row re-derived
against the 2.1.284 extraction. No existing verdict changed. Marker
drift was corrected where no verdict depended on it: `claude-api` ×3
gated, `design-sync` not hidden, and `skill-doctor` and `export` ×3
model-invocation-disabled. View regenerated (47 rows).
- **Rulings:** made 2026-09-29 by operator direction on the
orchestrator's recommendation. Each reason says so; **this review is the
human gate on them.**
- **`## Boundary` sections:** one per non-defer row, each with a
four-part reference file in the same skill. The sections are
presence-gated ("when the bundled `pr` skill resolves in this
session…"), name the native surface in a code span, and copy none of its
behavior.
- A **user-only** surface (`integration: suggest`) is offered to the
person: "you can run `/autofix-pr` instead of or alongside this".
  - A **model-invocable** surface gets a routing split.
- **Plugins touched** (patch bump and CHANGELOG entry each):
source-control, claude-config, verification, implementation, discovery,
claude-ops, context-budget, session-flow, planning, prototype,
debugging.
- `discovery:research-deep` no longer claims it can dispatch the bundled
`deep-research` workflow, because that workflow has model invocation
disabled. It now offers the workflow to the person.
- Frontmatter descriptions are **not** edited. Description phrases
follow one plugin per PR under the sweep contract in
`audit-native-overlap/SKILL.md`.

## Verification

- `overlap.py generate --check`: in sync (47 rows).
- `overlap.py self-check`: degraded, with the 2 documented advisories
only: older recorded versions in rows left for review, and no
`--upstream-sha`.
- `test_overlap.py`: 134 tests OK.
- `check-changed-skills.sh main`: 20 skills, 0 failed.
- `check-changelog-parity.sh`: `--check`, `--check-bump main`,
`--check-order` and `--check-preserved main` all exit 0.

## Related

- #5371, the inventory repair and dynamic overlap discovery this depends
on.
- **Left for a human ruling:** the two `design` rows (route) now extract
as `model_invocable: false`, which conflicts with `route`. The
extraction read one registration with an unresolved description, so the
evidence is uncertain. `plugin eval` and `morning` cannot be re-derived
from a binary extraction.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

---------

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
kyle-sexton added a commit that referenced this pull request Sep 29, 2026
…earning

Main now carries the squash merges of #5371 and #5387. Their files take
main's final form; this branch re-applies only its own hunks on top.

- overlap.py: the extracted build_native_index keeps main's plugin-backed
  reclassification; detect and dismiss both call it.
- test_overlap.py: main's PluginBackedSurfaceTests kept beside the
  dismissal tests.
- records.json: main's 47 rows plus this branch's 11 rows and 58
  dismissals. Two component fingerprints recomputed where main changed
  the description (architecture:map-context, claude-ops:audit-install-state);
  docs/native-surfaces.md regenerated.
- Versions above main: claude-ops 0.67.0, claude-config 0.53.2,
  source-control 0.62.29, native-references 3.2.2.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
kyle-sexton added a commit that referenced this pull request Sep 30, 2026
…e every candidate (#5466)

No related issue: operator decisions Q8, Q9 and Q12 from the 2026-09-29
native-surfaces interview (follow-up to #5387).

## Summary

Overlap discovery re-proposed the same false positives on every run,
because a ruling of "not an overlap" had nowhere to live. About 59% of
candidates were noise. 69 discovered candidates were also still unruled,
and the new inventory and detect behavior had no evals.

Stacked on #5387, which is stacked on #5371.

## Fix

- **Dismissals (Q8):**
- The store gains an optional `dismissals` list. Each entry records the
native surface, the component, a reason, `as_of` and date, plus a
fingerprint of each side's whitespace-collapsed description.
- New `overlap.py dismiss` subcommand. It refuses a pair that already
has a verdict row.
- `detect` suppresses a dismissed pair until either fingerprint changes,
then resurfaces it flagged `resurfaced: description changed`. A verdict
row always wins.
- `self-check` validates dismissals, and `generate` renders a Dismissed
table.
- **Triage (Q9):** all 69 remaining candidates are ruled.
- 11 verdict rows, 8 of them non-defer, each with a Boundary section, or
a registry row only for agents.
  - 58 dismissals.
  - Fresh detect: 0 new candidates, 58 suppressed, 0 resurfaced.
- Every reason ends "Ruled 2026-09-29 by operator direction on the
orchestrator's recommendation."
- **Evals (Q12):**
- inventory: `is-foo-real-under-a-degraded-lane`,
`docs-crosscheck-classifies-a-removed-command`
- audit-native-overlap: `user-only-native-recommends-suggest`,
`dismissed-pair-suppressed-then-resurfaced`
- **Version bumps:** claude-ops, source-control, bugs, github,
claude-memory, claude-config, each with a CHANGELOG entry.

## Verification

- `overlap.py generate --check`: in sync (58 rows, 58 dismissals).
- `overlap.py self-check`: degraded on the 2 documented advisories only.
- `test_overlap.py`: 156 tests OK (22 new). `overlap.test.sh`: exit 0.
`test_inventory.py`: 101 tests OK.
- Eval files validate against
`plugins/skill-quality/reference/evals.schema.json`, and
`check-evals-quality.sh` passes with 0 warnings. No model evals were
run.
- `check-changed-skills.sh main`: 25 skills, 0 failed.
- `check-changelog-parity.sh`: `--check`, `--check-order` and
`--check-preserved origin/main` pass. `--check-bump origin/main` is red
only on five plugins inherited from #5387, whose versions main has since
passed; they are renumbered when main merges into the stack.
- Pinned ruff: `check` and `format --check` clean.

## Related

- #5387 is the base, and #5371 below it.
- #5465 files future drift as work items; dismissals here are what keep
that intake from repeating false positives.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

---------

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
kyle-sexton added a commit that referenced this pull request Sep 30, 2026
No related issue: operator decisions Q5 and Q11 from the 2026-09-29
native-surfaces interview (follow-up to #5371).

## Summary

Nothing re-ran the inventory or the overlap detector when Claude Code
shipped a release. Drift was found only by hand: new surfaces, renamed
or removed ones, invocability changes, and decision rows whose recheck
trigger had fired.

Stacked on #5371; merge that first.

## Fix

- **`/claude-ops:changelog apply` Phase 7, native-surface drift**
(`context/native-drift.md`) runs:
  - `inventory.py --self-check`
  - a full `--binary-only --docs` extraction
  - `overlap.py detect` and `overlap.py self-check`
- `native_drift.py summarize`, then `diff` against the last good
summary. A broken extraction never becomes the baseline.
- **Report:** surfaces added, removed, renamed (by alias or description
similarity) and reclassified; invocability and marker changes; docs
cross-check changes; new overlap candidates; fired store triggers.
- **Work items are filed through `/work-items:track add`** (raw intake,
`needs-triage`), one per:
  - new candidate with no store row;
  - store row whose trigger fired;
- revalidation proposal: the CLI is past `VALIDATED_AGAINST` with every
lane ok and no surface change;
  - degraded or broken inventory.
- **Dedupe:** each item carries a
`native-drift:<kind>:<surface>:<component>` key. An open item with the
key is skipped, and a candidate closed as not planned counts as
dismissed.
- **Approval:** interactive runs confirm the batch once; unattended
runs, declared by the caller, file directly.
- **Dynamic:** no surface names are hard-coded.
- `claude-ops` 0.66.0.

## Verification

- `native_drift.test.sh`: 23 cases OK.
- `changelog-status.test.sh`: 70/70.
- `overlap.test.sh`: 134 OK. `test_inventory.py`: OK.
- `check-changed-skills.sh origin/main`: 5 skills, 0 failed. The three
warnings are on lines this PR does not touch.
- `check-changelog-parity.sh`: `--check`, `--check-order`, `--check-bump
origin/main` and `--check-preserved origin/main` all pass.
- Pinned ruff: `check` and `format --check` clean. markdownlint: 0
issues.
- **Live run on 2.1.285** (validated 2.1.284, every lane ok, identical
surfaces): the revalidate path. It also flagged six store rows whose
markers moved. Four are already corrected in #5387; the two `design`
rows move to `suggest` in the planned description sweep.

## Related

- #5371 is the base. #5387 corrects the four marker rows.
- The `native-drift` label is managed as code in github-iac; a follow-up
adds it. Until then items file without it, and dedupe does not depend on
the label.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

---------

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
kyle-sexton added a commit that referenced this pull request Sep 30, 2026
…ce lanes (#5467)

No related issue: operator decision Q10 from the 2026-09-29
native-surfaces interview (follow-up to #5371).

## Summary

The inventory covered commands, bundled skills and workflows, but not
built-in subagent types or built-in tools. Overlaps such as
`discovery:explore` vs the `Explore` agent, or `planning:plan` vs
`Plan`, were invisible to native-overlap detection.

Stacked on #5371; merge that first.

## Fix

- **`builtin_agents` lane:** 11 agents.
- **Found by shape:** an object with `agentType`, `source:"built-in"`,
and `whenToUse` or `getSystemPrompt`. Names come from literals or
resolved constants.
- **Roster status:** read from the roster function, as default,
conditional or absent.
  - **Canaries:** general-purpose, Explore, Plan, statusline-setup.
- **`builtin_tools` lane:** 80 named tools.
- **Found by shape:** `name` plus `maxResultSizeChars`. The builder's
minified name is never needed.
- **Names:** resolved from nearest-binding constants, preferring
PascalCase.
  - **Deferred marker:** where determinable.
- **Not counted as tools:** factory-built tools are counted separately,
and MCP templates are skipped.
  - **Canaries:** Bash, Read, Edit, Write, WebFetch.
- **Honest nulls:** descriptions built by functions stay null and are
counted under `integrity.undetermined`.
- **`--docs`:** adds a tools cross-check against the tools reference, as
its own block with its own status.
- **`overlap.py detect`:** scores both lanes as classes `builtin-agent`
and `builtin-tool`, `route` only. Each lane is presence-checked.
- **Native-references convention:** gains the rows for both classes
(3.2.0).
- **Speed:** `_nearest_binding` runs about 35x faster, so the whole run
stays at about 11.5 s.
- The four existing lanes are byte-identical to before.
- `claude-ops` 0.66.0.

## Verification

- `inventory.py --self-check`: all six lanes `ok`. It exits 3 (degraded)
only on the CLI version advisory, 2.1.285 against a validated 2.1.284.
- `test_inventory.py`: 128 tests OK (26 new).
- `test_overlap.py`: 137 tests OK. `overlap.test.sh` passes, and
`overlap.py generate --check` is in sync.
- `check-changed-skills.sh origin/main`: 17 skills, 0 failed.
- `check-changelog-parity.sh`: `--check`, `--check-order`, `--check-bump
origin/main` and `--check-preserved origin/main` all pass.
- Pinned ruff: `check` and `format --check` clean. typos: clean on the
changed files.
- **Detect on the new lanes:** 20 candidates. Top ones:
  - Plan → planning:plan (0.69)
  - Explore → discovery:explorer (0.60)
  - Explore → discovery:explore (0.57)

  They are ruled in a follow-up triage.

## Related

- #5371 is the base.
- #5466 has the dismissal mechanism the follow-up triage uses.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

---------

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
kyle-sexton added a commit that referenced this pull request Sep 30, 2026
Closes #5222

## Summary

Adds `/session-flow:tidy-work` (session-flow 0.41.0) and
`scripts/tidy_work.py`, an opt-in lifecycle for the gitignored `.work`
tiers in the current repo (the concern file's `memory_dir`, else
`.work`) and `~/.work`. Nothing runs unless the skill is invoked.

- `report` (read-only): per-item age, size, kind (handoff,
running-retro, checklist, slice, scratch, concern, unknown) and
in-flight status. A scratch item also shows the issue or PR its name
carries and that item's state.
- `normalize`: moves misplaced handoffs and running-retro ledgers (a
handoff's `.slots.json` sidecar with it) into the standard layout; never
deletes, never overwrites.
- `clean`: removes an item only when it is a known kind, is not in
flight, and names at least one issue or PR, every one closed or merged.
A scratch item goes only once its issue is closed or its PR merged.
- Unknown items (for example `.work/drain/status/*.json`) are always
kept and reported.

Issue body step 4 (entries that cannot be attributed are listed and
never deleted) holds for every kind. An item that names no issue or PR
is reported as `keep: names no issue or PR` and kept however old it is:
a handoff or running retro whose text names none, and every slice and
checklist, which have no attribution source (no writer records an issue
or PR in them), so `clean` never removes one. The owner comment's
`clean` (Expected 3) is read as acting on the attributed items that are
not in flight.

A handoff or running-retro file is that kind wherever it sits (the root,
`handoffs/`, `running-retros/`), so `report` and `normalize` agree on
it, with its `.slots.json` sidecar part of it.

Attribution (issue body steps 2 to 4, owner comment's `scratch` kind):

- A top-level entry is `scratch` when its name holds exactly one
all-digit token of 3 to 7 digits, optionally prefixed `pr`, `issue`, or
`gh` (`lint-5371.log`, `measure-4608`, `pr4120.md`, `reverify-4186`,
`scratch-4586-d2cc1ea4d`). A year-like token (1900 to 2099) needs the
prefix: `pr2026.md` is attributed, `backup-2026.tar` and `notes-2025.md`
are not. A name with no such token, with several (a version such as
`native-surfaces-2-1-284`, a date such as `triage-2026-09-28`), or
starting with a writer's `<TS>Z-` timestamp is not attributed: it stays
unknown and is never removed.
- The number is read as an issue or PR of the repository holding the
memory root, and its state is looked up: `gh api` lists the open items
once per repository, then reads each other number once. A number that is
no issue or PR of the repository is unknown, so it keeps the item. A PR
closed without merging is `closed-unmerged` and keeps the item.
- `report` and each `clean` dry-run path show every state looked up, for
example `stale [#4608 closed]` or `would remove: <path> [#5371 merged]`,
so the confirmation covers the issue or PR each path was matched to
(handoffs and retros included).
- The issue lists the name, a marker file, or the producing branch as
attribution sources. Only the name is read for scratch: the repo defines
no marker-file convention, and no `.work` writer records a branch
(handoff frontmatter is
`type`/`handoff_shape`/`date`/`topic`/`session_id`/`transcript`/`previous_handoff`/`chain`).
Step 3 offers delete or consolidate; `clean` deletes.
- In `~/.work` a bare number has no repository, so a scratch item there
is unknown-state and always kept.

In flight, so kept:

- a slice whose `INDEX.md` `status:` is not `done`, or that holds a
child slice whose status is not `done` (the topic-docs contract makes
that frontmatter field the slice's state; missing and unrecognized
values keep it too)
- a workflow checklist with an unfinished stage
- anything with a `.git` file or directory under it: a clone or worktree
can hold commits that exist nowhere else, and the tracked-path guard
only asks the outer repository
- a change inside `--days` (default 14)
- a later handoff that is itself kept and names the item. A stale
handoff does not keep what it names.
- a handoff or running retro whose text names an issue or PR that is not
closed or merged (a `github.com` URL, `owner/repo#N`, or `#N`), or one
whose state could not be read
- a scratch item whose number is open, a PR closed unmerged, or
unreadable

No writer records an issue or PR in frontmatter, so handoff and retro
references are read from the text. A bare `#N` means the repository
holding the memory root; in `~/.work` it counts as unknown. A link is
looked up only for an item nothing cheaper already keeps, except a
scratch item's, whose state is the attribution the report shows. The
`.git` check likewise runs only on an item nothing cheaper keeps.

The contract designates no top-level `.work` directory as scratch by
name, so the entries of `reviews/`, `exports/`, `overengineering/`,
`enforceability/`, `docs-hygiene/`, and `lanes/` are kind `concern`:
reported, never removed, because their owning skills read them back.
Other tools' free-form output without an issue or PR number (for example
`~/.work/local-otel-claude-code`) is unknown and kept.

## Fix

Placement is session-flow, not repo-hygiene as the issue title says: the
issue body proposes `/session-flow:tidy-work`, session-flow owns the
`.work` layout and the `save_point.py` frontmatter parsers, repo-hygiene
cannot import them without a cross-plugin dependency, and `~/.work` is
not a repo. The "call disk-hygiene's protection list" line is a proposed
direction, not an expected item, and disk-hygiene has no public entry
point for it (only its skill-private `baseline-policy.json`), so reading
it would break skill encapsulation. It is not read.

`normalize` and `clean` are dry runs that print exact absolute paths
until `--apply`. They never modify content git tracks, following the
topic-docs runtime guards:

- a memory root whose `.gitignore` lacks a line `*` is refused
(`~/.work` excepted, which has no fixed self-ignore file)
- every command, `report` included, rejects an existing memory root
outside the repository whose `.gitignore` lacks a line `*` (exit 2,
nothing listed). `memory_dir` comes from the repo-controlled
`.claude/topic-docs.yaml` and `report` is pre-approved, so `memory_dir:
/etc` or `../outside` is refused before anything is enumerated. This is
the guard the handoff writer requires of every memory root; a root below
the repository is unaffected, and a root that does not exist yet reports
as empty
- an item with a tracked path under it is refused
- every command rejects a memory root that is the repository root
(`--memory-dir .`)
- a relative `--memory-dir` resolves against the repository top level
- `--apply` also refuses paths outside the resolved roots and symlinks
leaving them

The skill resolves `memory_dir` through `parse-concern-value.sh` as
handoff does and passes `--memory-dir`. It pre-approves only the
read-only `report` invocation and the read-only parse helper; `--apply`
stays behind the permission flow. Version bump 0.40.2 to 0.41.0 (minor),
changelog entry above main's 0.40.x entries, README, `topic-docs.md`,
catalog and cheat sheet updated.

## Verification

- `uvx pytest -q plugins/session-flow/scripts/tests/test_tidy_work.py`:
45 passed (186 with `test_save_point.py`). Nine cases fail against the
previous `tidy_work.py`: the unattributed handoff, running retro, `done`
slice and finished checklist are kept and absent from the dry run and
`--apply` (tree listing unchanged); a kept handoff that names no issue
keeps the scratch directory it names; the year names (`_name_refs`
directly, and `backup-2026.tar` and `notes-2025.md` beside a closed
`#2026` and a merged `#2025` in the link table) stay unknown while
`pr2026.md` is attributed and removed once `#2026` is closed. The
symlink-escape, tracked-content and missing-self-ignore cases now remove
an attributed scratch directory or handoff (`measure-4608` with `#4608`
merged, a handoff naming a closed `#7`) instead of an unattributed one,
and each fails when its guard is disabled. The stale-handoff-chain case
runs on attributed handoffs and a scratch directory they name. The
memory-root cases still cover an outside root (absolute and
`../outside`) rejected by all three commands with nothing listed and
reported once its `.gitignore` holds `*`, an absent outside root, every
file `normalize` would move classified as a known kind, and `clean
--apply` on a stale root handoff with its sidecar. The scratch cases use
real `.work` names: six attribute (`lint-5371.log`, `measure-4608`,
`pr4120.md`, `reverify-4186`, `scratch-4586-d2cc1ea4d`, `pr2026.md`) and
eleven do not; dry-run and apply remove only closed or merged scratch; a
number that is no issue or PR keeps its item; an item holding a `.git`
directory or file survives even with a merged number.
- `scripts/run-ruff.sh check plugins/session-flow/scripts`: all checks
passed; `format --check` clean on `tidy_work.py` and its test.
- Read-only `tidy_work.py report --days 0 --memory-dir <this repo's
.work>` with live `gh` (`--days 0` turns off the recency window; at the
default 14 days every entry there is kept): 87 items, 6 removable, 81
kept. Removable: `lint-5371.log` `[#5371 merged]`, `measure-4608`
`[#4608 closed]`, `pr4120.md` `[#4120 closed]`, `pr5425-body.md` `[#5425
merged]`, `scratch-4586-d2cc1ea4d` `[#4586 closed]`, and one handoff
whose text names only merged and closed items, with its sidecar.
`reverify-4186` is kept because it holds five git clones and #4186 is
open; `native-surfaces-2-1-284` and `triage-2026-09-28` are unknown and
kept; each of the 14 handoffs names an issue or PR, so none is kept for
lacking one. `clean --days 0` without `--apply` listed those paths with
their `[#N state]`; `--apply` was not run.
- `scripts/check-changelog-parity.sh --check --check-order`: pass;
`scripts/validate-plugins.sh`: pass; `generate-catalog.mjs --check` and
`generate-cheatsheet.mjs --check`: in sync
- `CHECK_SKILL_SKILLS_ROOT=plugins/session-flow/skills bash
plugins/skill-quality/scripts/check-skill.sh tidy-work`: PASS, 0
warnings; `check-evals-quality.sh` on the skill's `evals.json`: 0
warnings; `check-purged-em-dashes.sh`: no em dashes; `markdownlint-cli2`
on the touched markdown: 0 issues
- `bash scripts/affected-tests.sh --run --jobs 4` (shell suites, rerun
on this head): all pass except `scripts/check-script-contract.test.sh`,
whose two `check-html-assets.sh` cases need `npm ci` (htmlhint absent),
unrelated to this change. The selection also names two Python suites the
shell runner does not execute, `test_tidy_work.py` (run above with `uvx
pytest`) and `test_audit_skill_visibility.py` (untouched by this branch)
- `git ls-files -s plugins/session-flow/scripts/tidy_work.py`: mode
100755, as the lint exec-bit step requires of a shebang file
(`save_point.py` is 100755)
- `bash plugins/session-flow/scripts/tidy_work.test.sh`: SKIP locally
(system python3 lacks pytest); the pytest run above covers it

## Related

Related: #3555, #4295, #4228.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

---------

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant