fix(claude-ops): keep native drift report-only without a store and validate its inputs - #5525
Conversation
overlap.py self-check exits 3 in report-only mode when the store is absent; Phase 7 read every exit 3 as passing, so it filed tracker items and replaced the baseline. native_drift.py diff now marks a run whose --store names no file report_only, moves its items to unfiled with no overflow, and the phase keeps the previous baseline. native_drift.py also validates each loaded input's shape (summary, detect, store, inventory) and exits 2 on JSON that parses but has the wrong shape, instead of a traceback with exit 1. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…ding parse_tools_table stopped only at a non-table line after a row matched. When the table's rows stopped matching (names without backticks) it read on through the page and took a later backticked row as the table, so the tools block reported every tool undocumented instead of broken. The scan now ends at the next heading whether or not a row matched. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Codex Review SummaryThis comment shows the latest Codex review activity on this pull request.
ℹ️ About Codex in GitHubYour team has set up Codex to review pull requests in this repo. Reviews are triggered when you
Codex reacts with 👀 while any review is running, comments if it has suggestions, and reacts with 👍 once all reviews finish with no findings. |
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: eed6e03450
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
Security review and Codex findings on the report-only change:
- shape_error checks leaf types (aliases, markers, descriptions, store
row native name/class, summary candidates), requires the store's
schema 1, and requires the detect report's schema, discovery,
integrity and candidates, so `{}` can no longer become a candidate
baseline. A TypeError or AttributeError left over while processing
loaded input exits 2 instead of a traceback.
- load maps RecursionError (deeply nested JSON) and ValueError (an
integer literal past the digit limit) to exit 2.
- clip collapses whitespace and replaces backticks, and each item gets
a quote block that puts every fact in a code span on a "> " line, so
upstream text cannot forge a Drift key line and suppress a filing,
open a fence, or mention a user or issue.
- diff's report_only is a required keyword, and filing also needs the
overlap self-check to have found a valid store.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Reviewing PR #5525
|
|
Claude finished @kyle-sexton's task in 2m 9s —— View job Security review — no issues found
Reviewed the full diff ( What this PR touches, security-relevant angle by angle:
No CRITICAL, IMPORTANT, or SUGGESTION findings to report. |
|
@codex review |
…drift markdownlint MD038 failed on the `> ` code span. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: 3eeea28940
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
|
Claude finished @kyle-sexton's task in 2m 39s —— View job Claude finished the security review
Security reviewReviewed the diff ( No security findings. Specifically:
No command injection, path traversal, SSRF, or credential-handling issues in the changed files; |
|
Claude finished @kyle-sexton's task in 4m 1s —— View job Reviewing PR #5525 (
|
Codex re-review on #5525: a schema-1 detect candidate with empty `native` and `component` objects passed and keyed the baseline as None:None:None, so the next valid run could file the real candidate as new. native.name, component.plugin and component.skill must now be strings; a real detect report still validates. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…rt-only # Conflicts: # plugins/claude-ops/.claude-plugin/plugin.json # plugins/claude-ops/CHANGELOG.md
|
Claude finished @kyle-sexton's task in 1m 55s —— View job Security review — no issues found
Note on scope: the dispatched HEAD SHA ( Claude's prior pass on this PR (commit Codex P2 — "Reject detect reports with missing required fields": closed. Codex P2 — "Validate nested summary records before diffing": closed. Re-checked the rest of the security surface for regressions from these two commits:
No CRITICAL, IMPORTANT, or SUGGESTION findings to report. |
|
Claude finished @kyle-sexton's task in 2m 14s —— View job Claude finished the security review
Security reviewScope: No security findings.
|
|
Claude finished @kyle-sexton's task in 4m 18s —— View job Reviewing PR #5525 (
|
…esolved (#5619) No related issue: operator-directed fix from the 2.1.284/2.1.285 native-surface review (ClaudeDesign had an empty description, so detect never paired it). ## Summary `inventory.py` left 14 built-in descriptions unresolved on Claude Code 2.1.285 (13 tools plus the `design` bundled skill). It now resolves 13 of them. `design` stays unresolved because its text reads a table keyed by a runtime mode. Detect also scores native names that have no description, and native drift now files each newly unresolved description. ## Fix - **Inventory (`inventory.py`).** A call with arguments (`kbr(RTe())`, `gLr(void 0)`, `P({...})`) is now followed into its function, with that function's parameters shadowed. Template substitutions resolve, and so do `||`/`??` fallbacks (an empty `""` fallback is skipped), parenthesized parts such as `d+(x()?m:c)+p`, and `[...].join(sep)` arrays. A template made only of runtime parts stays unresolved. - **Module-scoped identifiers.** The 2.1.285 bundle concatenates about 2,100 modules, and minified names repeat between them. An imported name resolves to the one top-level declaration in the one module that exports it. Any other name resolves inside its own module. A name that is neither imported nor declared in its module is unresolved. This fixes a wrong value that would otherwise appear: `workflow-authoring`'s `${jd}` reads `Workflow`, not another module's local `host_exit`. A single-letter function resolves only when it is the one top-level declaration in its module. - **Detect (`discover.py`).** PascalCase native names are split into words (`ClaudeDesign` scores as "design", `EnterWorktree` as "enter worktree"). `user_facing_name` is scored. It is not added to the dismissal fingerprint, so `audit-native-overlap/SKILL.md` stays accurate. - **Drift intake (`native_drift.py`).** `summarize` records `integrity.undetermined.description_unresolved`. `diff` adds an `unresolved-description` item (key `native-drift:unresolved-description:<name>:inventory`) for each name the previous summary did not list. These items go through the existing key-based filing path, with no label. `context/native-drift.md` documents the new kind and its title. - `claude-ops` 0.75.1 -> 0.76.0, with a CHANGELOG entry. `docs/native-surfaces/records.json` and all skill bodies are unchanged. ## Verification - `test_inventory.py`: 144 tests OK (10 new synthetic-bundle cases, negative cases included). `test_overlap.py`: 182 OK. `test_native_drift.py`: 40 OK. `overlap.test.sh` and `native_drift.test.sh` exit 0. The pinned `scripts/run-ruff.sh check`/`format`, `markdownlint-cli2` and `typos` all pass. - `inventory.py --self-check` on the installed 2.1.285 exits 3 with only the version advisory, the same result as `origin/main`. All lanes are ok. The run takes about 12 s (11 s before). - Old and new extractions were compared on 2.1.284 and on 2.1.285. Unresolved descriptions drop 18->2 on 2.1.284 and 14->1 on 2.1.285. Every other changed field was checked by hand against the bundle, for example Grep `ALWAYS use Grep ... as a Bash command`, `workflow-authoring` `Workflow`, and Edit user-facing name `Update`. - `overlap.py detect --inventory <2.1.285> --repo .` goes from 23 to 50 candidates. The operator rules on the new candidates and on the dismissals that came back because their descriptions changed. ## Related - #5465, #5525, #5574 (native-drift intake this extends) 🤖 Generated with [Claude Code](https://claude.com/claude-code) --------- Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>

No related issue: follow-up to late review findings on the merged #5465 (Codex P1 and P2) and #5467 (Claude correctness).
Summary
/claude-ops:changelog applyPhase 7 treated everyoverlap.py self-checkexit 3 as passing. With no overlap store, self-check exits 3 and declares the run report-only, but the phase still filed work items and promoted its summary to the baseline. That could happen in a foreign repository, or after the store went missing.native_drift.pyaccepted any syntactically valid JSON. A wrong-shaped input, such as a[]summary, crashed with a traceback and exit 1 instead of the documented exit-2 input error.parse_tools_tabledid not bound its scan. If the docs page reformatted the tools table, the parser absorbed unrelated backticked rows from later sections. It then reported every real tool asundocumentedinstead of reporting itselfbroken.Fix
native_drift.py diffsets"report_only": truewhen--storeis omitted or names no file, the same testoverlap.pyself-check uses. Items move tounfiled, nothing is filed, and the baseline is kept.context/native-drift.mdbranches on that field, not on the exit code.load()checks each input kind's top-level and nested shape (summary,detect,store,inventory) and exits 2 withmalformed <kind> input <path>: ….parse_tools_tablestops at the first heading after the table header, whether or not any row matched.Verification
native_drift.test.sh: 33 tests OK. New tests cover report-only filing nothing, a present store filing normally, and wrong-shaped inputs for every flag; they fail on the old code.test_inventory.py: 133 tests OK. The reviewer's reproduction now yields{}and abrokenblock.changelog-status.test.sh: 70/70.check-changed-skills.sh origin/main: 0 failed.check-spoke-plugin-root.sh,validate-plugin-contracts.mjs(0 warnings) andgenerate-catalog.mjs --check: all pass.check-changelog-parity.sh: all four modes pass. typos: clean.Related
🤖 Generated with Claude Code