Repository navigation
Proactively suggest an independent second-harness review flow - #643
Conversation
Pure, repo-aware aggregation behind the forthcoming list_reviewer_vendors MCP tool: given the daemons the server reports and the current session's repo identity, compute the reviewer vendors that can actually run an unattended review flow for this repo (intersection of repo-hosting daemons and their unattended vendors, deduped, conservative model-override AND). Every empty result is disambiguated by a single typed reason with a fixed precedence, so it never reads as a cause it is not. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Adds a read-only list_reviewer_vendors tool to `kcap mcp flows` that reads GET /api/daemons and reports, for the current session's repo, which vendors can actually run an unattended review flow right now — so a caller can offer a reviewer that will not be rejected. Tolerant daemon parsing feeds the pure aggregation; the driver harness is inferred from its env markers (claude, codex; anything ambiguous or unverified stays null) and echoed as driver_vendor. Result DTO is source-generated (AOT-clean). Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
The model-driven skill that offers an independent review flow at the two milestones (spec finalized, implementation complete), recommends a reviewer via list_reviewer_vendors, and hands off to review-flows only on explicit consent. Ships identically to every harness (driver vendor is inferred at call time, not stamped). Registered in SourceNames + help-plugin.txt + README (all test-pinned), with conformance pins on the triggers and the safety guardrails (never auto-run, ordinary self-review stays local, unknown driver never claims a different model). Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
The reproducible acceptance harness for the model-driven skill's trigger surface, kept out of the shipped skill dir. Documents the two-layer gate (CI-durable conformance pins + deterministic unit tests; dev-time skill-creator eval on the reference harness), the precommitted thresholds, the full positive/negative/availability corpus, and the known automated-multi-harness-eval gap. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
PR Summary by QodoProactively suggest a second-harness review flow with repo-aware reviewer lookup
AI Description
Diagram
High-Level Assessment
Files changed (14)
|
Code Review by Qodo
1.
|
list_reviewer_vendors made the GET (and its persistent-401 return) before Aggregate, so an unresolved repo plus an unauthorized daemons lookup surfaced an auth error instead of the contractual repo_unresolved result. repo_unresolved is a local precondition and the highest-precedence reason, so short-circuit it before any network call — which also spares a pointless request. Test pins that an unresolved repo returns repo_unresolved and never hits the server. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
- Security: repo.identity now carries a non-sensitive owner/repo id, never the absolute local repoRoot (which stays internal, used only for the on-disk hosting match) — no local path leaks to the model. - Reliability: the read-only tool now reads the machine id via MachineId.ReadPersisted() instead of Get(), so it never persists machine.json (a null id simply drops the same-machine filter for that call). - Robustness: repo-path matching unifies separators and folds case per-OS (OrdinalIgnoreCase on Windows/macOS, Ordinal on Linux), so a '\'-vs-'/' or trailing-slash difference no longer yields a false no_repo_hosting_daemon. - Standard: ParseDaemons uses JsonElementExtensions helpers instead of raw JsonValueKind checks (repo compliance rule). Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Local testing against a real daemon exposed that the server serializes /api/daemons with JsonNamingPolicy.SnakeCaseLower, so ParseDaemons was looking up camelCase field names that never matched — every daemon record was skipped and the tool always returned schema_skew (i.e. the feature never worked in production). Parse the snake_case names (repo_paths, machine_id, unattended_vendors, unattended_vendor_capabilities, supports_reviewer_model_resolution) and update the unit/WireMock fixtures to the real wire. Verified end to end: the tool now returns the daemon's eight unattended reviewers for a hosted repo. Also update the integration tools-list assertion (8 -> 9) and pin list_reviewer_vendors in the advertised set — the CI break from adding the tool. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Linear: AI-2175
What & why
Most vibe coders never run an independent code review — not with a second coding agent, nor outside the implementing session. Review flows already exist and the reviewer can be any installed harness independent of the driver, but the
review-flowsskill is strictly opt-in and never proposes itself, so beginners never discover it.This adds a model-driven skill that proactively offers an independent review flow at the two natural milestones — right after a spec/design is finalized, and right after implementation is complete — and recommends a reviewer harness that will actually run for this repo.
Changes (purely additive — one new skill + one new tool)
list_reviewer_vendorsMCP tool inkcap mcp flows(agent-only, read-only). ReadsGET /api/daemonsand computes, client-side, the reviewer vendors that can run an unattended review for this repo now — the intersection of same-machine daemons hosting the repo × their unattended vendors. Every empty result is disambiguated by a single typedreason(repo_unresolved · schema_skew · lookup_failed · no_daemons_connected · no_repo_hosting_daemon · no_unattended_reviewer). Conservativemodel_override(AND across daemons). Tolerant daemon parsing (case-insensitive; malformed records skipped + counted; all-unparseable ⇒ schema_skew).suggest-review-flowskill — offers at spec-complete →spec-reviewand implementation-complete →code-review; prefers a reviewer ≠ the driver, falls back to a same-vendor independent flow, and maps an empty result to one-line unlock guidance. Never auto-runs — it offers, and hands off toreview-flowsonly on explicit consent. Ships identically to all nine harnesses; the driver vendor is inferred at call time from harness env (not stamped, since six harnesses share~/.agents/skills), echoed asdriver_vendor, unknown ⇒ null ⇒ no "different model" claim.DriverVendor— env-marker inference (claude, codex today; ambiguous/unverified ⇒ null).AgentsSkillsInstaller.SourceNames+help-plugin.txt+README.md(all test-pinned); conformance pins on the triggers and safety guardrails.docs/eval/suggest-review-flow.md— the reproducible triggering-eval corpus + methodology.No server change
Verified against kcap-server:
GET /api/daemonsalready returnsDaemonInfowithrepoPaths+unattendedVendors+machineId+unattendedVendorCapabilities, so this is CLI-only.Review
The design (AI-2175) was hardened over a Codex spec-review flow, 6 rounds to clean (repo-affinity, a circular eval gate, and a branch-name identity collision were all caught there). Deterministic unit/contract tests cover the tool's aggregation, every typed reason + precedence, partial-record handling, and identity edge cases; static conformance pins carry the skill's model-driven surface. AOT publish is clean.
No submodule pin bump is included (that's the maintainer's).
🤖 Generated with Claude Code