Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
2 changes: 1 addition & 1 deletion plugins/review/.claude-plugin/plugin.json
Original file line number Diff line number Diff line change
@@ -1,7 +1,7 @@
{
"$schema": "https://json.schemastore.org/claude-code-plugin-manifest.json",
"name": "review",
"version": "0.15.1",
"version": "0.15.2",
"description": "Code-review toolkit: six read-only reviewer agents (code, security, architecture, doc drift, build/test/lint, CI-log audit) plus two orchestration skills — a single-lens quality gate and a multi-surface review fan-out with severity-ranked, deduplicated findings.",
"author": {
"name": "Melodic Software",
Expand Down
25 changes: 25 additions & 0 deletions plugins/review/CHANGELOG.md
Original file line number Diff line number Diff line change
Expand Up @@ -3,6 +3,31 @@
All notable changes to the `review` plugin are documented here. Format follows
[Keep a Changelog](https://keepachangelog.com/en/1.1.0/); this plugin uses semantic versioning.

## [0.15.2]

### Fixed

- Carried the `code-review` framing reconciliation of `0.15.1` into the `fanout` skill, which
described the same nonexistent plugin independently: `SKILL.md`'s "Orchestrator plugins" section
and `context/findings-normalization.md` both listed `code-review` as one of three optional
`claude-plugins-official` orchestrator plugins invoked as `/code-review:code-review`. `SKILL.md`
now carries its own "Boundary" section for the two real surfaces, and
`findings-normalization.md`'s per-surface parse-contracts table no longer lists `code-review` as
a normalized fan-out leaf. Two `fanout` evals (`pr-comment-gate-opt-in`, renamed
`unscored-surface-severity-derived-not-invented`) carried the same stale framing and were updated
for internal consistency. The two surface descriptions are not restated — `SKILL.md` points at
`pr.md`'s Boundary for those and carries only the fan-out-specific reasoning.
- `fanout`'s exclusion of the bundled command no longer rests on classing a **bare**
`/code-review` invocation as PR-mutating. Per <https://code.claude.com/docs/en/code-review>
("Review a diff locally"), bare `/code-review` is report-only — findings arrive in the
conversation, and only `--fix` and `--comment` mutate — matching the gate scoping `pr.md` already
applies. The Boundary section now states the real reason it is not a normalized leaf (it is
itself a multi-agent review of the same diff, with no documented output schema to write a parse
contract against) and points the reader at running it directly (review-caught).
- The README's optional-orchestrator roster also names `codex` (OpenAI Codex marketplace), the
other orchestrator `fanout` dispatches, and points at both skills' Boundary sections
(review-caught).

## [0.15.1]

### Fixed
Expand Down
13 changes: 7 additions & 6 deletions plugins/review/README.md
Original file line number Diff line number Diff line change
Expand Up @@ -58,12 +58,13 @@ Invoke via `@review:<agent>` or let Claude delegate.
files when present (the marketplace-wide ecosystem-commands contract,
`docs/conventions/ecosystem-commands/README.md`), falling back to your documented conventions,
then its own bundled generic defaults as a last resort.
- **Graceful degrade.** The optional `pr-review-toolkit` orchestrator plugin (from the official
marketplace) adds adversarial breadth when installed; every path works without it. Claude
Code's bundled `/code-review` command and the managed Code Review GitHub App service are
separate built-in/managed surfaces, not marketplace plugins — see the Boundary section of
[`skills/quality-gate/context/pr.md`](skills/quality-gate/context/pr.md) for how the `pr` mode
relates to them.
- **Graceful degrade.** Optional orchestrator plugins — `pr-review-toolkit` from the official
marketplace and `codex` from the OpenAI Codex marketplace — add adversarial breadth when
installed; every path works without them. Claude Code's bundled `/code-review` command and the
managed Code Review GitHub App service are separate built-in/managed surfaces, not marketplace
plugins — see the Boundary sections of
[`skills/quality-gate/context/pr.md`](skills/quality-gate/context/pr.md) and
[`skills/fanout/SKILL.md`](skills/fanout/SKILL.md) for how each skill relates to them.
- **Self-contained.** The severity baseline and all mode guidance ship inside the plugin
and are referenced via `${CLAUDE_PLUGIN_ROOT}`.

Expand Down
9 changes: 7 additions & 2 deletions plugins/review/skills/fanout/SKILL.md
Original file line number Diff line number Diff line change
Expand Up @@ -83,12 +83,17 @@ Run the self-ignore guard ("Shared inputs"), then write the ranked report to `<f

## Orchestrator plugins

Three optional orchestrator plugins add adversarial breadth — two same-vendor Claude plugins from the `claude-plugins-official` marketplace, plus the OpenAI Codex plugin (`codex@openai-codex`) as a different-model surface. All run on the MAIN THREAD (they fan out their own agents; a subagent cannot dependably do that). Each is a graceful enhancement, not a hard dependency:
Two optional orchestrator plugins add adversarial breadth — `pr-review-toolkit` from the `claude-plugins-official` marketplace, plus the OpenAI Codex plugin (`codex@openai-codex`) as a different-model surface. Both run on the MAIN THREAD (they fan out their own agents; a subagent cannot dependably do that). Each is a graceful enhancement, not a hard dependency:
Comment thread
kyle-sexton marked this conversation as resolved.

- **`pr-review-toolkit`** — `/pr-review-toolkit:review-pr`: aspect-scoped agent fan-out. Absent → this plugin's leaf agents cover most of the same dimensions; note that orchestrator breadth was skipped.
- **`code-review`** — `/code-review:code-review`: parallel reviewers + confidence scorer for an existing PR. **PR-mutation gate:** its PR mode posts findings as a PR comment, which violates the review modes' report-only contract; when the branch has an open PR, dispatch it only on explicit user opt-in ("post the review comment"), otherwise skip it and name the skip in `## Surfaces`. Absent → note the skip; a repository's own CI review bot (when present) still provides PR coverage.
- **`codex`** (OpenAI Codex) — `/codex:review`: read-only cross-vendor review, so it satisfies the review modes' report-only contract with no PR-mutation gate; `/codex:adversarial-review`: red-teams the diff, fitting the intentional adversarial-breadth intent. The first surface backed by a **different model** — its blind spots are uncorrelated with the same-vendor leaf agents and Claude orchestrators, so a finding only Codex raises is signal the rest structurally cannot see. Invoke it with `--wait` so the review runs in the foreground and returns findings in the same turn (its default prompts and may run in a background task the synchronous normalization step would miss), and pass `--base <review-base>` carrying this skill's resolved review diff base ("Shared inputs") so Codex diffs the same change set as every other dispatched surface — without it Codex auto-picks the working tree or default branch. Absent → note the skip; the same-vendor surfaces still cover most dimensions.

## Boundary — built-in/managed surfaces, not marketplace plugins

Two Claude Code surfaces overlap this skill's job on an open PR: the **bundled `/code-review` command** and the **managed Code Review GitHub App service**. Neither is an installable `claude-plugins-official` marketplace plugin — both ship with Claude Code itself. [`quality-gate/context/pr.md`](../quality-gate/context/pr.md)'s "Boundary" section describes what each one is and what it mutates; that description is not repeated here. What is specific to this skill:

Neither is dispatched as a normalized fan-out surface and [context/findings-normalization.md](context/findings-normalization.md) carries no parse contract for either — but for different reasons. The managed service posts its findings to the PR instead of returning them to this skill. Bare `/code-review` **is** report-only (findings arrive in the conversation; only `--fix` and `--comment` mutate); it is left out because it is itself a multi-agent review of the same diff, overlapping this skill's own leaf reviewers, and its finding output has no documented schema to write a parse contract against. Run it directly when you want that second orchestration — this skill names the overlap in `## Surfaces` rather than dispatching it. **PR-mutation gate:** `/code-review --comment` and triggering the managed service both post to the PR, which violates the review modes' report-only contract; when the branch has an open PR, invoke either only on explicit user opt-in ("post the review comment"), otherwise note the overlap in `## Surfaces` without invoking it. Not enabled/available → note the skip; a repository's own CI review bot (e.g. the managed service, when enabled) still provides PR coverage independently.

## What this skill does NOT do

- **Review modes do not apply fixes** — mutation happens only through the explicit `fix` action; it never auto-runs after a review.
Expand Down
Original file line number Diff line number Diff line change
Expand Up @@ -13,11 +13,12 @@ The 5-stage main-thread pipeline that turns heterogeneous free-text findings fro
| `architecture-guardian` | Violations / Risks / Opportunities | — | file-only (Violations); none (Risks/Opportunities) |
| `doc-drift-detector` | Stale / Missing / Aspirational | — | doc-file line (table) |
| slice-subagents | project's tiers (or baseline) | — | `file:line` (inferred) |
| `code-review` plugin | none (flat issue list) | 0–100, filters <80 | GitHub permalink `#L[s]-L[e]` |
| `pr-review-toolkit` orchestrator | Critical / Important / Suggestion | — | `[file:line]` (inferred) |

Line numbers from LLM reviewers drift — treat inferred lines as approximate and keep dedup noise-tolerant.

**Not in this table:** the bundled `/code-review` command and the managed Code Review GitHub App service (SKILL.md "Boundary — built-in/managed surfaces, not marketplace plugins"). The managed service posts its findings to the PR rather than returning them to normalize; bare `/code-review` is report-only, but is itself a multi-agent review of the same diff whose output has no documented schema to parse. Neither is dispatched as a fan-out leaf here.

## Stage 0 — Extraction (Sonnet)

Per-surface free-text → records `{surface, file, line, line_basis, category, native_severity, native_confidence, raw_text}`.
Expand All @@ -34,7 +35,7 @@ Map native severity → the tier vocabulary in effect (the project's own, else `
- code-reviewer, slice-subagents, pr-review-toolkit: identity mapping (Critical/Important-or-Warning/Suggestion).
- architecture-guardian: Violation → CRITICAL (broken rule today) or IMPORTANT (drift) by content; **Risk → SUGGESTION + `forward-flag: future` (NEVER a blocking tier)**; Opportunity → SUGGESTION.
- doc-drift: Stale → IMPORTANT; Missing/Aspirational → SUGGESTION.
- **Surfaces emitting no severity (e.g. the `code-review` plugin)** → DERIVE from content: bug/correctness → CRITICAL or IMPORTANT by impact; convention-adherence → IMPORTANT; ambiguous → IMPORTANT + `pending: human-tier`. A confidence filter having passed is confidence-of-realness, NOT severity — a high-confidence nitpick is still a nitpick.
- **Surfaces emitting no severity** → DERIVE from content: bug/correctness → CRITICAL or IMPORTANT by impact; convention-adherence → IMPORTANT; ambiguous → IMPORTANT + `pending: human-tier`. A confidence filter having passed is confidence-of-realness, NOT severity — a high-confidence nitpick is still a nitpick.

## Stage 2 — Confidence enum (deterministic / Haiku)

Expand Down
8 changes: 4 additions & 4 deletions plugins/review/skills/fanout/evals/evals.json
Original file line number Diff line number Diff line change
Expand Up @@ -65,10 +65,10 @@
"id": 6,
"name": "pr-comment-gate-opt-in",
"prompt": "Fan out a full review on this branch — it has an open PR.",
"expected_output": "The skill withholds the code-review orchestrator's PR mode (which posts findings as a PR comment and would violate the report-only contract) unless the user explicitly opts in, skipping it and naming the skip in the surfaces list instead.",
"expected_output": "The skill withholds `/code-review --comment` and the managed Code Review GitHub App service (neither dispatched as a report-only fan-out surface, since both post findings as PR comments and would violate the report-only contract) unless the user explicitly opts in, naming the skip in the surfaces list instead.",
"files": [],
"expectations": [
"Output does NOT dispatch the code-review orchestrator's PR-comment-posting mode without explicit user opt-in",
"Output does NOT invoke `/code-review --comment` or trigger the managed Code Review service without explicit user opt-in",
"Output names the skipped PR-mutation surface in its surfaces list rather than silently omitting it",
"Output keeps the review report-only, posting nothing to the open PR absent an explicit opt-in"
]
Expand Down Expand Up @@ -124,8 +124,8 @@
},
{
"id": 11,
"name": "plugin-findings-severity-derived-not-invented",
"prompt": "[Scenario: a PR-present change where the code-review plugin emits confidence-filtered issues.] Fan out a review where the code-review orchestrator plugin returns confidence-filtered issues with no native severity.",
"name": "unscored-surface-severity-derived-not-invented",
"prompt": "[Scenario: a dispatched surface returns confidence-filtered issues with no native severity label.] Fan out a review where one dispatched surface returns confidence-filtered issues with no native severity.",
"expected_output": "The severity crosswalk derives a tier from finding content for surfaces with no native severity; a high confidence score is never treated as a severity tier.",
"files": [],
"expectations": [
Expand Down
Loading