feat(source-control): babysit autopilot merge tier, gate-enforced and disabled by default - #665
Conversation
… disabled by default Add the #476 autopilot merge tier to babysit-prs: a second bot account (author != approver) runs a genuine review pass and submits an approving review only when clean, after which the pinned merge gate merges only when every criterion holds. Criteria are enforced deterministically in babysit_merge.py behind the fail-closed --autopilot-merge-tier umbrella flag: - required checks green incl. the review workflow (mergeStateStatus CLEAN; ruleset never bypassed) and head pinned by --expected-head as always; - issue-linked (a closing-issue reference); - authored by a configured pipeline lane (--lane-logins); - no human CHANGES_REQUESTED / blocking comment / unresolved review thread; - no configured do-not-merge label (--block-labels); - a distinct-bot approving review on the live head (--approver-bot-logins; author != approver, head SHA unchanged since review). Any criterion failing falls back to today's behavior: the PR is reported on the human merge-ready list. The umbrella flag refuses (exit 3) unless all three parameter sets are supplied, and every criterion predicate is reused from the shared babysit_classify module rather than re-implemented. Absent the flag the merge gate is byte-for-byte its prior self, so worker/autopilot's existing gate-proven merges are unchanged. The tier exists only while the new babysit_autopilot_merge_tier userConfig (boolean, default off) is enabled; enabling it and any later gate-off flip is a separate, announced operator step. safety.md's "Never do automatically: merge" contract is updated deliberately to codify the tier and its criteria. Tests: new test_babysit_merge.py covers each criterion with a passing and a fall-back fixture plus the tier-absent no-network invariant; guard and skill-contract tests cover the fail-closed CLI and the tier prose. Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: d775a4d86c
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
…rge-tier # Conflicts: # plugins/source-control/.claude-plugin/plugin.json # plugins/source-control/CHANGELOG.md
|
Claude finished @kyle-sexton's task in 3m 21s —— View job PR Review
|
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: a9268072f6
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
…uality cap The #476 tier subsection duplicated the criteria that reference/safety.md already codifies and pushed SKILL.md to 530 lines, over the skill-quality hard cap of 500. Condense the subsection to an operational summary (disabled by default, the review-pass-then-gated-merge flow, fail-closed, fallback) and point to safety.md as the single home for the criteria and the safety-contract rationale, removing the duplication. Retarget the skill-contract test at the condensed prose and add a safety.md guard for the criteria that moved there. Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
|
Claude finished @kyle-sexton's task in 3m 40s —— View job PR Review
New commit:
|
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: da239d82c7
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
…nd approval checks Address three P1 review findings on the #476 tier, all fail-closed: - Ignore superseded bot approvals: collapse reviews to each actor's latest decisive state (via babysit_feedback.latest_reviews_by_author, decisive_only=True) before accepting a tier approval, so an approver bot that approved then requested changes on the same head no longer counts even when a separate approval keeps the base reviewDecision APPROVED. - Recognize explicit do-not-merge comments: has_blocking_text carries no do-not-merge pattern, so a plain human "please do not merge" veto escaped the "no human blocking comment" criterion. Add a bounded merge-veto regex to the tier's human-comment scan. - Scan inline review comments: a human inline veto on a since-resolved thread saw neither the base unresolved-thread gate nor the tier scan. Include the paginated pull-request review comments (fetch_pull_request_review_comments) with normalized author data in the human-blocking corpus. Each fix reuses an existing shared seam; new fixtures cover all three. Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
|
Claude finished @kyle-sexton's task in 3m 4s —— View job PR Review
New commit:
|
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: fddd27cf50
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
… approver bot is_bot() accepts any [bot]/Bot-typed login before consulting approver_bot_logins, so an unrelated installed GitHub App's APPROVED review satisfied the tier's author != approver criterion, defeating the babysit_approver_bot_logins boundary. find_distinct_bot_approval now also requires normalized membership in the configured approver set, so only the designated bot's approval authorizes a tier merge. Fail-closed: an empty approver set matches nothing (already refused by the umbrella flag). Test: test_unconfigured_bot_approval_is_rejected. Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
|
Claude finished @kyle-sexton's task in 1m 39s —— View job PR Review
New commit:
|
| Finding | Status |
|---|---|
| Superseded bot approval not detected | RESOLVED (fddd27c) |
| Explicit do-not-merge comment not caught | RESOLVED (fddd27c) |
| Inline review comments excluded from human-veto corpus | RESOLVED (fddd27c) |
| Unconfigured bot approval satisfies tier | RESOLVED (9912da4) |
One minor item remains: Finding 2 — test_label_match_is_case_insensitive missing self.assertFalse(result["ready"]). This is a one-line addition with no risk. The do-not-merge hold is intentional and the operator sign-off path is correct. The PR is substantively clean pending that small test fix.
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: 9912da47b7
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
…sit merge tier The triage lane records a defaulted (maintainer-vetoable) decision only as a 'Decision defaulted: X — veto before merge' issue comment, which a deterministic merge gate cannot see (Codex P1 on #641). Add a tier criterion: the PR's linked issue must carry no unratified decision-default marker. A marker is ratified only by a human maintainer (authorAssociation OWNER/MEMBER) comment strictly after it; reactions are not consulted (the reactions API carries no author association, so a reaction cannot be attributed to a maintainer, and using the operator's self-logins would let pipeline automation clear its own veto — the same #450 attribution-drift hazard). Fail closed: an unratified marker, or a linked issue whose comments cannot be fetched, holds the PR for the human merge-ready list. Marker matching is deliberately loose (over-matching only holds more for the human). Reuses fetch_issue_comments and the shared actor_kind predicate; criterion codified in safety.md. Pass, fall-back, and fetch-error fixtures added. Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
|
Claude finished @kyle-sexton's task in 4m 8s —— View job PR Review
New commit:
|
| Finding | Status |
|---|---|
| Superseded bot approval not detected | RESOLVED (fddd27c) |
| Explicit do-not-merge comment not caught | RESOLVED (fddd27c) |
| Inline review comments excluded from human-veto corpus | RESOLVED (fddd27c) |
| Unconfigured bot approval satisfies tier | RESOLVED (9912da4) |
| Decision-default veto criterion missing | RESOLVED (a12e770) |
test_label_match_is_case_insensitive missing assertFalse(result["ready"]) |
Open — minor |
One minor test-completeness item remains. The do-not-merge hold is intentional; the operator sign-off path is correct.
…rge-tier # Conflicts: # plugins/source-control/.claude-plugin/plugin.json # plugins/source-control/CHANGELOG.md
|
Claude finished @kyle-sexton's task in 1m 51s —— View job PR Review
New commits:
|
| Precondition | Status |
|---|---|
| Decision-default veto added as a merge criterion | ✅ In this branch (a12e770) |
#642 fix (count_classified fail-open) landing |
✅ Now in this branch (9d7ca0c via d8b561d) |
Both preconditions are now satisfied in the branch. Tier activation remains a separate, announced operator step; the config flip is not part of this PR. The do-not-merge hold and required operator sign-off are still correct and intentional.
Finding 2 — Still open (minor): test_label_match_is_case_insensitive missing ready=False
test_babysit_merge.py L206–210
Carried across six commits now without being addressed. The test asserts blockingLabels == ["do-not-merge"] but never asserts result["ready"] is False. A refactor that populated blockingLabels without pushing the matching entry into blockers would pass this test silently. Every adjacent blocking test — test_blocking_label_blocks, test_non_lane_author_blocks, test_unratified_decision_default_marker_blocks — makes the assertFalse(result["ready"]) assertion; this test should too.
Summary
The merge with origin/main is a clean administrative sync — no conflicts, no code changes to the tier. Its only significance is that the #642 fix is now in the branch, completing the second of two stated flip preconditions.
Full status across the review cycle:
| Finding | Status |
|---|---|
| Superseded bot approval not detected | RESOLVED (fddd27c) |
| Explicit do-not-merge comment not caught | RESOLVED (fddd27c) |
| Inline review comments excluded from human-veto corpus | RESOLVED (fddd27c) |
| Unconfigured bot approval satisfies tier | RESOLVED (9912da4) |
| Decision-default veto criterion missing | RESOLVED (a12e770) |
test_label_match_is_case_insensitive missing assertFalse(result["ready"]) |
Open — minor |
The implementation is substantively clean. The one remaining item is a one-line test assertion that should be added before merge, but it is a test-completeness gap — not a correctness defect in the gate itself. The do-not-merge hold is intentional and the operator sign-off path is correct.
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: d8b561de5f
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
… the decision-default veto The decision-default veto treated ANY later OWNER/MEMBER comment as ratifying the marker, so an unrelated "thanks" would clear the maintainer veto (Codex P1 on #665). Ratification now requires an explicit signal in the maintainer's comment: a closed, whole-word token set (ratify/ratified, approve/approved, confirm/confirmed), and not a withheld-approval negation (reusing the shared NON_APPROVAL_RE). Matching is strict/fail-closed — an unrelated comment, a signal that predates the marker, or a negated approval do not clear, and only over-hold to the human list. Token set documented in safety.md next to the marker contract. Fixtures added: explicit-signal clears (existing), unrelated-comment does not, signal-before-marker does not, negated-approval does not. Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: d3c9bbf38b
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
…rge-tier # Conflicts: # plugins/source-control/.claude-plugin/plugin.json # plugins/source-control/CHANGELOG.md
|
Claude finished @kyle-sexton's task in 3m 50s —— View job PR Review
New commits:
|
| Finding | Status |
|---|---|
| Superseded bot approval not detected | RESOLVED (fddd27c) |
| Explicit do-not-merge comment not caught | RESOLVED (fddd27c) |
| Inline review comments excluded from human-veto corpus | RESOLVED (fddd27c) |
| Unconfigured bot approval satisfies tier | RESOLVED (9912da4) |
| Decision-default veto criterion missing | RESOLVED (a12e770) |
| Any later maintainer comment ratifies the veto (implicit ratification) | RESOLVED (d3c9bbf) |
| Cross-repo issue fetch uses wrong repo | Open |
test_label_match_is_case_insensitive missing assertFalse(result["ready"]) |
Open — minor |
The do-not-merge hold is intentional; the operator sign-off path is correct. Two items remain before merge: the cross-repo issue fetch gap (correctness, rare failure path) and the one-line test assertion gap (test completeness only, not a gate defect).
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: ec9fc5b5b2
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
…erenced issue's own repo The decision-default veto fetched the linked issue's comments from the PR's repository using only the issue number, so a PR closing an issue in ANOTHER repo would read a same-numbered issue in the wrong repo and miss its marker (Codex P1 on #665). Each `closingIssuesReferences` entry carries its own repository; the scan now reads from `owner/name` (falling back to the PR repo), and the held list reports canonical `owner/repo#n` refs. Fixture added: test_cross_repo_linked_issue_read_from_its_own_repo. Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
… in the tier veto and ratification scans An approver configured via `--approver-bot-logins` (or a pipeline lane login) that GitHub reports as account type `User` without a `[bot]` suffix passed `find_distinct_bot_approval` — which classifies via `is_bot`, honoring the configured login over the non-`Bot` typename — but was read as HUMAN by the tier's veto and ratification scans, which classify via `actor_kind`. Root cause: `actor_kind` returned "human" on a `User` typename / `is_bot: False` before ever consulting `extra_bot_logins`, so threading the config in was provably inert (a populated config still returned "human"). That defeats the field's own documented purpose — "accounts whose metadata misreports them as users" — and disagrees with `is_bot`, which already ignores any non-`Bot` typename. Consequence: the configured approver's own clean review body could count as a human blocking comment against the very approval it provides, and a configured automation identity could ratify a decision-default marker — automation clearing its own veto, the #450 attribution-drift hazard the ratification design explicitly excludes (Codex P2 on #665, thread PRRT_kwDOTCGFQM6SOgye). Fix, fail-closed: `actor_kind` now lets an explicitly configured `extra_bot_logins` identity outrank a `User` typename, matching `is_bot` and honoring the field's contract; default-config callers are unchanged. The tier threads a single automation config (`approver_bot_logins | lane_logins`) into both the human-blocking scan and the decision-default ratification scan via a new `AutopilotMergeTierConfig.automation_actor_config` property, so a configured identity is always automation for those scans — never a human veto, never a ratifier. Fixtures added: classifier-level precedence contract; a configured `User`-type approver's blocking-looking review body does not block; a configured `User`-type lane account with an OWNER association does not ratify a decision default; unconfigured humans unchanged. Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
|
Claude finished @kyle-sexton's task in 2m 16s —— View job PR Review
New commits:
|
| Finding | Status |
|---|---|
| Superseded bot approval not detected | RESOLVED (fddd27c) |
| Explicit do-not-merge comment not caught | RESOLVED (fddd27c) |
| Inline review comments excluded from human-veto corpus | RESOLVED (fddd27c) |
| Unconfigured bot approval satisfies tier | RESOLVED (9912da4) |
| Decision-default veto criterion missing | RESOLVED (a12e770) |
| Any later maintainer comment ratifies the veto (implicit ratification) | RESOLVED (d3c9bbf) |
| Cross-repo issue fetch reads wrong repo | RESOLVED (a9bb74d) |
| Later maintainer revocation of a ratified default not honored | RESOLVED (871a9e8) |
| Configured approver/lane login misclassified in veto and ratification scans | RESOLVED (c404522) |
| Codex P1: review-workflow status not independently gated | Deferred to #675 flip precondition |
test_label_match_is_case_insensitive missing assertFalse(result["ready"]) |
Open — minor |
Nine correctness findings resolved across the review cycle, each with dedicated test cases. The decision-default veto is now fully hardened end-to-end: explicit signal required, revocations honored, cross-repo issues read from their own repo, and automation identities can never be confused for human ratifiers regardless of GitHub's reported typename. One minor test-completeness gap remains. The do-not-merge hold is intentional; the operator sign-off path is correct.
|
Operator sign-off recorded (2026-07-20, in-session confirmation to control tower): safety digest verified all hardening claims (per-claim table in tower session log; digest read inline at tower after the auto-mode classifier correctly blocked delegating this review). Operator confirmed MERGE via explicit in-session prompt. Behavioral inertness at merge: tier engages only when |
|
Babysit worker pass (safe tier, merge-gate ON — no merge/approve/label change). Classification of the open item from the latest review run (
Disposition of the rest of the corpus:
New head: |
…702) ## Summary Dispatched workers assigned an out-of-tree per-issue worktree can silently operate from the wrong checkout: the shell's working directory does not reliably persist across separate tool calls, so a one-time `cd` at session start does not survive into later commands. In #566 this caused an implementation worker's edits to land in the canonical checkout instead of its assigned worktree — caught only because that worker happened to notice. The fix hardens the text that briefs future dispatched workers so they don't repeat the failure. ## Fix `implement-dispatch`'s "Compose the brief" step (`plugins/implementation/skills/implement-dispatch/SKILL.md`) — the single owner of worker-brief composition mechanics, which `work` Step 5 and any plan-routed dispatch already delegate to — now requires a worktree-cwd clause whenever a worker edits in a dedicated worktree: the brief MUST give the worktree's absolute path and instruct the worker to never rely on cwd persisting across separate tool calls, anchoring every file-touching command with `git -C <worktree-path>` (or a re-`cd` per call) rather than a one-time `cd`. Stated as imperative discipline (the load-bearing mitigation), not as an asserted universal tool defect — the mitigation is correct regardless of whether the persistence gap reproduces everywhere. `implementation` bumped 0.7.3 → 0.7.4 with a matching CHANGELOG entry. **Deferred (parked):** `source-control`'s `babysit-prs` Worker Contract (`reference/orchestration.md`) is a second lane that hands a worker a worktree path and has the same latent bug. Its manifest is already being bumped by open PR #665 (currently CONFLICTING with main), so this PR does not race that `plugin.json`. The one-line babysit-prs mitigation should be applied as a follow-up once #665 lands (or folded into #665). ## Verification ``` $ bash scripts/check-changelog-parity.sh --check-bump origin/main Every plugin whose version changed vs origin/main has a '## [<version>]' CHANGELOG.md entry. $ npx markdownlint-cli2 --config .markdownlint-cli2.jsonc \ plugins/implementation/skills/implement-dispatch/SKILL.md plugins/implementation/CHANGELOG.md Summary: 0 error(s) $ typos --config _typos.toml plugins/implementation/ (exit 0) $ node scripts/validate-plugin-contracts.mjs Plugin contracts validated: 33 setup skills and 1811 plugin files checked. ``` Closes #566 ## Related - #665 — bumps `source-control`'s manifest; blocks bundling the parallel babysit-prs fix here (see Fix). - #572 — deferred autonomous branch/worktree provisioning seam across `work`, `implement-dispatch`, `worktree`. --------- Co-authored-by: Claude <noreply@anthropic.com>
…rge-tier # Conflicts: # plugins/source-control/.claude-plugin/plugin.json # plugins/source-control/CHANGELOG.md
|
Claude finished @kyle-sexton's task in 1m 32s —— View job PR Review
New commits:
|
| Finding | Status |
|---|---|
| Superseded bot approval not detected | RESOLVED (fddd27c) |
| Explicit do-not-merge comment not caught | RESOLVED (fddd27c) |
| Inline review comments excluded from human-veto corpus | RESOLVED (fddd27c) |
| Unconfigured bot approval satisfies tier | RESOLVED (9912da4) |
| Decision-default veto criterion missing | RESOLVED (a12e770) |
| Any later maintainer comment ratifies the veto (implicit ratification) | RESOLVED (d3c9bbf) |
| Cross-repo issue fetch reads wrong repo | RESOLVED (a9bb74d) |
| Later maintainer revocation of a ratified default not honored | RESOLVED (871a9e8) |
| Configured approver/lane login misclassified in veto and ratification scans | RESOLVED (c404522) |
test_label_match_is_case_insensitive missing assertFalse(result["ready"]) |
RESOLVED (ab9d9e5) |
| Codex P1: review-workflow status not independently gated | Deferred to #675 flip precondition |
All ten correctness/test findings are now resolved. The review cycle is complete. The do-not-merge hold is intentional; operator sign-off was recorded by kyle-sexton (2026-07-20). Tier activation, #675, and the loud announcement remain as the separate gate-off flip steps.
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: 88ee348756
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
…ng findings The #476 autopilot merge tier accepted a distinct-bot APPROVED review on identity/head alone. The human-blocking corpus scan deliberately skips bot-authored items and GitHub can still return reviewDecision=APPROVED, so an approver bot that approved while its own body reported a P1/CRITICAL/blocking regression sailed through the gate -- a fail-open in a merge gate. Severity-scan the accepted approval's own body: a structured high-severity finding (CRITICAL/IMPORTANT surviving negation redaction, or a P0-P3 severity badge / bracketed [P0-P3] marker) invalidates it as a tier approval. Treated as no approval, not as a human blocker, and a since-superseded earlier clean approval is not honored past the latest blocking verdict. Only the shared classifier's structured severity vocabulary is scanned -- prose words ("blocking"/"regression") are not, so a clean approval describing the fix it signs off stays accepted. This is the autopilot-tier answer to the open #621 question of whether formal APPROVED-state reviews should be severity-scanned.
|
Claude finished @kyle-sexton's task in 2m 56s —— View job PR Review
New commit:
|
| Test | Path exercised | Correct |
|---|---|---|
test_approval_with_shields_p1_badge_body_rejected |
SEVERITY_BADGE_RE (/badge/P1-orange) |
✅ |
test_approval_with_bracketed_p1_body_rejected |
SEVERITY_PLAIN_RE ([P1]) |
✅ |
test_approval_with_critical_body_rejected |
has_blocking_severity (CRITICAL) |
✅ |
test_approval_with_important_body_rejected |
has_blocking_severity (IMPORTANT) |
✅ |
test_clean_approval_body_accepted |
clean body → accepted | ✅ |
test_negated_severity_approval_body_accepted |
negation redaction clears CRITICAL/IMPORTANT |
✅ |
test_lowercase_severity_prose_approval_accepted |
lowercase prose doesn't match structured vocab | ✅ |
The _assert_rejected helper correctly enforces the full invariant: distinctBotApproval is None, ready=False, humanBlockingComments == [] (not a human blocker), and a "blocking findings in its body" message present. This is the right set of assertions — it verifies both the absence of the approval and the correct blocker category.
The predicate-level ApprovalReportsBlockingUnit class adds independent unit coverage of approval_reports_blocking(body) — confirming CRITICAL, IMPORTANT, [P0], [P1], and shields badge all return True, while empty body, LGTM, negated CRITICAL, and lowercase prose all return False. Fix is correct and complete.
Complete finding status across the review cycle
| Finding | Status |
|---|---|
| Superseded bot approval not detected | RESOLVED (fddd27c) |
| Explicit do-not-merge comment not caught | RESOLVED (fddd27c) |
| Inline review comments excluded from human-veto corpus | RESOLVED (fddd27c) |
| Unconfigured bot approval satisfies tier | RESOLVED (9912da4) |
| Decision-default veto criterion missing | RESOLVED (a12e770) |
| Any later maintainer comment ratifies the veto (implicit ratification) | RESOLVED (d3c9bbf) |
| Cross-repo issue fetch reads wrong repo | RESOLVED (a9bb74d) |
| Later maintainer revocation of a ratified default not honored | RESOLVED (871a9e8) |
| Configured approver/lane login misclassified in veto and ratification scans | RESOLVED (c404522) |
test_label_match_is_case_insensitive missing assertFalse(result["ready"]) |
RESOLVED (ab9d9e5) |
| Blocking findings in approver-bot review body accepted as clean approval | RESOLVED (b3bedc9) |
| Codex P1: review-workflow status not independently gated | Deferred to #675 flip precondition |
All eleven correctness/test findings are now resolved. The do-not-merge hold is intentional; operator sign-off was recorded by kyle-sexton (2026-07-20). Tier activation, #675, and the loud announcement remain as the separate gate-off flip steps. The review cycle is complete.
…, not just be required (#675) Codex P1 (re-review of 8e780ea) on #706: the review-workflow requiredness precondition over-claimed. It said making the review workflow a required status context "closes that hole deterministically", but requiredness alone does not guarantee the review RAN. A conditionally-skipped review job can report a SKIPPED conclusion that is counted as a passing state — this skill's own babysit_checks.py classifies SKIPPED as success (CHECK_SUCCESS_STATES) — so a required-but-skipped review still reads mergeStateStatus == CLEAN without having gated anything, and the enabled tier would proceed on its bot approval alone. Strengthen the operator precondition (docs-only, in #706's own scope): - Enabling now requires the review workflow to be a required status context AND to always run to a non-skipped conclusion on every PR to the protected base. - Qualify "deterministically": requiring the context closes the hole only when the workflow cannot conditionally skip on the tier's PRs. - Extend "do not enable the tier" to the skippable-on-those-PRs case. Code-side enforcement (a merge-gate predicate that verifies the review context actually executed, or exempting the review context from SKIPPED-as-success) belongs in the feature (#476/#665), not this docs PR, and is left for that surface. Refs #675 Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
…, not just be required (#675) Codex P1 (re-review of 8e780ea) on #706: the review-workflow requiredness precondition over-claimed. It said making the review workflow a required status context "closes that hole deterministically", but requiredness alone does not guarantee the review RAN. A conditionally-skipped review job can report a SKIPPED conclusion that is counted as a passing state — this skill's own babysit_checks.py classifies SKIPPED as success (CHECK_SUCCESS_STATES) — so a required-but-skipped review still reads mergeStateStatus == CLEAN without having gated anything, and the enabled tier would proceed on its bot approval alone. Strengthen the operator precondition (docs-only, in #706's own scope): - Enabling now requires the review workflow to be a required status context AND to always run to a non-skipped conclusion on every PR to the protected base. - Qualify "deterministically": requiring the context closes the hole only when the workflow cannot conditionally skip on the tier's PRs. - Extend "do not enable the tier" to the skippable-on-those-PRs case. Code-side enforcement (a merge-gate predicate that verifies the review context actually executed, or exempting the review context from SKIPPED-as-success) belongs in the feature (#476/#665), not this docs PR, and is left for that surface. Refs #675 Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
…ecify the approve mechanic and review-context precondition (#675) (#706) Closes #675 **Gate-off flip precondition** for the autopilot merge tier (#476). All changes are prose/contract + test-pinning: **no behavioral change** to the merge gate or any script, and the tier still ships **DISABLED**. Implements the operator-approved RECOMMENDED options recorded on #675: - **§3 wiring — fork (b), reference-file restructure.** Autopilot's step 3 in `SKILL.md` no longer inlines a base-only merge command that would ignore the tier flags when the tier is enabled. It now points at `reference/safety.md`, the single home for both the base and the enabled-tier merge paths, so an ENABLED config can no longer merge via the flagless base path. Chosen over raising the skill-quality line cap (`SKILL.md` sits at 499/500). - **Second-account approve mechanic (specified).** The concrete out-of-band approval the gate's distinct-bot criterion requires: `gh pr review … --approve` under a distinct `<approver-bot-logins>` identity (`GH_TOKEN` or `gh auth switch`, never the PR author or a lane identity), only after a genuine clean review pass, on the live head so `--expected-head` holds. - **Review-context requiredness — fork (a), required-context precondition.** Enabling the tier now carries a documented operator precondition: the base branch ruleset must make the review workflow a **required** status context, so `mergeStateStatus == CLEAN` actually proves the review ran. Chosen over a merge-gate review-context config (rejected fork b) to keep the gate deterministic with nothing new to wire — and deliberately kept distinct from the existing snapshot-side `babysit_review_gate_context` (a review-trigger timing signal, not a merge predicate). The skill-contract tests are split/extended to pin all three contracts against drift (the worker-tier push paragraph, which has no merge tier, keeps its inline-command assertion). ## Related - Stacked on #665 (`feat/476-autopilot-merge-tier`) and **must land after it** — this PR is based on that branch, not `main`; retarget to `main` once #665 merges. - **Version reconciliation at rebase:** bumped to `0.14.1`, continuous with the current base (`feat/476` @ `0.14.0`). #665's plugin version is being re-composed to `0.15.0` (collision with #660's `0.14.0`); when #665 lands, rebase this and re-bump to one increment past whatever `main` then carries (expected `0.15.1`). - Precedent for the flip-precondition designation: #642, #476. ### Verification - `engine.test.sh`: 312 tests pass, ruff clean, guarded-wrapper behavior checks pass. - `check-skill.sh babysit-prs`: PASS (SKILL.md 499/500, internal refs resolve, markdownlint clean). - markdownlint on `safety.md` + `CHANGELOG.md`: 0 errors. 🤖 Generated with [Claude Code](https://claude.com/claude-code) --------- Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
…path (#1299) Closes #704 ## Summary `source-control:babysit-prs` handed a dispatched worker a worktree path the same way `implementation:implement-dispatch` did before #566/#702, and with the same latent correctness bug: nothing in the Worker Contract or the Worker Prompt Template told the worker that the shell's working directory does not reliably persist across separate tool calls. A worker that `cd`s into its assigned worktree once can drift back to the session's default checkout on a later call and land a branch-owned fix in the wrong repository — the exact failure #566 hit. #702 fixed the sibling lane and explicitly parked this one, because `source-control`'s manifest was being bumped by the then-open #665. **#665 merged on 2026-07-20**, so the version-file race is gone and this lands the parked mitigation. ## Fix `plugins/source-control/skills/babysit-prs/reference/orchestration.md`: - **Worker Contract, dispatch inputs** — the orchestrator must hand over the worktree's **absolute** path, never a relative one (a relative path resolves against whatever cwd happens to be on the call that uses it). - **Worker Contract, worker obligations** — a new rule: never rely on cwd persisting across separate tool calls, decomposed into the three classes of call that need anchoring (below). - **Worker Prompt Template** — the same instruction in the text a worker actually receives, plus `Worktree (absolute path):` in the field block. The Contract list is what the orchestrator reads; the template is what reaches the worker, so both need it. **Where this diverges from a straight mirror of #702.** `git -C` is not a blanket answer in this lane. Babysit workers also read and edit files by path, and run commands that have no `-C` and derive their target from the working directory. So the obligation names three classes: | Class | Anchor | |---|---| | git — `status`, `add`, `commit`, `diff`, `log`, `push` | `git -C <absolute-worktree-path>` | | file read / edit / write / glob / search | an absolute `<absolute-worktree-path>/…` path, never a bare relative one | | bare `gh` | `GH_REPO=owner/repo` — documented by `gh help environment` for every subcommand that "otherwise operate[s] on a local repository" | | `scripts/fetch-all-pr-comments.sh` | `FETCH_COMMENTS_OWNER` / `FETCH_COMMENTS_REPO` (the override the skill already documents) | | the target repo's own build/test/lint | re-`cd` per call | The filesystem row and the `GH_REPO` form both came out of review — see the threaded replies on `370059bc`. The filesystem class is the sharper of the two: a relative path handed to a read tool resolves against cwd exactly as a shell command does, so a worker could validate a finding against the session's checkout (silently producing a wrong classification) while its `git -C` calls correctly targeted the assigned worktree. The contract also states which helper is *not* affected, so a worker does not needlessly wrap it: `source-control-babysit-resolve-thread` takes the PR and every thread pin as explicit arguments and reads nothing from the working directory (verified — no cwd or `gh repo view` derivation in `babysit_resolve_thread.py`; the wrapper's own `cd` only resolves its script directory). `source-control` 0.26.0 → 0.26.1 with a matching CHANGELOG entry. Per the [plugins reference](https://code.claude.com/docs/en/plugins-reference) (fetched this session), a `plugin.json` with an explicit `version` only delivers changes to installed consumers when that field is bumped — pushing commits alone leaves the cached copy in place. ## Deliberately not in `SKILL.md` #702 landed in two surfaces of `implement-dispatch`: the contract **and** a `## Gotchas` reminder. Only the contract half fits here. `babysit-prs/SKILL.md` is **499 lines** against `plugins/skill-quality/scripts/check-skill.sh`'s `LINE_HARD_CAP=500`, so the always-loaded surface has zero budget — the Gotchas bullet was written, pushed the file to 507 lines, failed `scripts/check-changed-skills.sh`, and was reverted. That is not a loss of reach: the Worker Contract and its prompt template are mandatory reading for any orchestrator composing a brief, so the mitigation reaches every worker without a third copy to keep in sync. Note the rule grew during review — three anchored classes rather than one — which makes the third copy more expensive to maintain, not less, and sharpens the case for keeping it at one owner. The general constraint is filed as #1297. ## Test plan The behavior this PR mitigates was observed first-party in this session — the harness itself reported it while running a backgrounded command from the worktree: ```text Session cwd remains D:\repos\github.com\melodic-software\claude-code-plugins; directory changes made by the backgrounded command do not apply to subsequent commands. ``` Stated in the diff as imperative discipline rather than an asserted universal tool defect (same framing #702 chose): the mitigation is correct regardless of where the persistence gap reproduces. ```text $ bash scripts/check-changelog-parity.sh --check-bump origin/main Every plugin whose version changed vs origin/main has a '## [<version>]' CHANGELOG.md entry. $ node scripts/validate-plugin-contracts.mjs Plugin contracts validated: 43 setup skills and 2079 plugin files checked. $ npx markdownlint-cli2 --config .markdownlint-cli2.jsonc \ plugins/source-control/skills/babysit-prs/reference/orchestration.md \ plugins/source-control/CHANGELOG.md Summary: 0 error(s) $ typos --config _typos.toml plugins/source-control/ # exit 0 $ claude plugin validate . # Validation passed $ scripts/check-changed-skills.sh origin/main # exit 0, no findings $ scripts/check-skill-portability.sh origin/main # exit 0 $ python -m unittest discover -s tests -t . # plugins/source-control/skills/babysit-prs/scripts Ran 348 tests in 33.490s — OK $ bash plugins/source-control/skills/babysit-prs/scripts/engine.test.sh OK · ruff: All checks passed · 4/4 guarded-wrapper behavior PASS ``` ### Review rounds Pre-PR: an independent fresh-context reviewer (`review:code-reviewer`) read the full current files rather than the diff alone. Verdict sound, no CRITICAL/IMPORTANT. Its four SUGGESTIONs landed in `dc1bd3e2` — the exempt-helper justification narrowed to the one helper a worker actually invokes, the concrete escapes inlined into the prompt template so a worker acting purely off the pasted text has them, and the absolute-path rationale widened from "relative to the orchestrator's checkout" to any relative path. Post-PR, three follow-up commits: - `62aaa2e5` — terminal period on the new obligation bullet, and the template's verbatim `text` block reflowed to break at clause boundaries rather than mid-phrase. - `370059bc` — the filesystem-path anchoring class, and the `gh` escape moved to `GH_REPO`. One finding was classified **INCORRECT on its premise but adopted on its conclusion**: two reviewers independently read `gh -R owner/repo` as an unrunnable flag placement. It is in fact runnable (`gh -R <repo> pr view 1299` → exit 0 on gh 2.95.0), but two independent misreadings are evidence a dispatched worker could misread it too, so the doc moved to `GH_REPO` — which has no placement to get wrong and covers subcommands that define no `--repo` flag at all. Their suggested `gh pr <subcommand> -R owner/repo` wording was not taken: it hardcodes the `pr` command group into a rule that also governs `gh api` and `gh run`. ## Related - #566, #702 — the sibling `implement-dispatch` fix this mirrors; #702's own body parked this lane. - #665 — bumped `source-control`'s manifest and blocked bundling this into #702. Merged 2026-07-20. - #1297 — filed from this work: `babysit-prs/SKILL.md` sits at 499/500 lines, so the next contract change to this skill cannot reach its always-loaded surface either. - #1300 — filed from this PR's review, VALID but deferred: `reference/review-discipline.md`'s mandatory finding-extractor subagent prompt has the same gap one dispatch level down. Deferred because that file is the plugin-scope seam `/source-control:pull-request` also consumes, which is outside #704's brief. 🤖 Generated with [Claude Code](https://claude.com/claude-code) https://claude.ai/code/session_01NE8W4XPPVfZVCGBW2GYgyH --------- Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>

Summary
Implements the #476 autopilot merge tier for
babysit-prs: at day-scale throughput, humanapprove-and-merge is the pipeline bottleneck. The tier lets the fleet satisfy the branch
ruleset instead of bypassing it — a second bot account (author ≠ approver) runs a genuine review
pass and submits an approving review only when clean, after which the pinned merge gate merges
only when every criterion holds. The ruleset itself is never touched; the bot review is what
makes this a real gate rather than a rubber stamp.
HELD FOR OPERATOR: do-not-merge until the operator personally reviews the safety-contract
change. This PR moves through the normal pipeline as a proposal only. The babysit safety contract
(
reference/safety.md) changes if and only if the operator merges this PR. It carries thedo-not-mergelabel at creation;do-not-merge.ymlwill hold the required check red by design —that is the intended hold, not a CI failure to fix.
Criteria (verbatim from the 2026-07-19 15:49 maintainer decision)
Enforced deterministically in
babysit_merge.pybehind the fail-closed--autopilot-merge-tierumbrella flag. Base-gate criteria already existed; the tier layers the rest.
mergeStateStatusCLEAN + required-context reconciliation)closingIssuesReferencesnon-empty--lane-loginsmembershipreviewDecision)has_blocking_text/has_blocking_severityover human comments--block-labelsclosingIssuesReferencescomments for aDecision defaultedmarker; ratified only by a humanOWNER/MEMBERcomment after it--expected-headas always--approver-bot-logins, sharedis_bot)No deviations from the decision comment. Every criterion predicate is reused from the shared
babysit_classifymodule (#634), not re-implemented. Any criterion failing falls back to today'sbehavior: the PR is reported on the human merge-ready list — the tier never routes around the gate.
Flip preconditions (tier stays disabled until these land)
Both are enforcement-in-the-gate placements recorded on #476; the pre-tier world (a human reads the
comments at GATE-ON) stays safe meanwhile.
deliberately not consulted for ratification: the reactions API carries no
authorAssociation, so areaction cannot be attributed to a maintainer, and attributing it via the operator's self-logins
would let pipeline automation clear its own veto (the source-control:babysit-prs: gh-bot.sh bot token generation fails, worker writes fall back to personal identity #450 attribution-drift hazard). Ratification
is therefore a maintainer comment, or a manual merge. Note: feat(source-control): babysit autopilot merge tier — bot-account review + criteria-gated merge #476's own thread quotes the marker
phrase, so once the tier is active feat(source-control): babysit autopilot merge tier, gate-enforced and disabled by default #665 would self-hold under this criterion — harmless, since it
is
do-not-merge-held and merged by the operator by hand regardless.count_classifiedfail-open (stale pipe-rows in non-thread comments) has landed(source-control 0.13.3 via fix(source-control): scope babysit classification credit to its surface (#642) #666) and is merged into this branch, so this flip-precondition is now
cleared. It was never worked around locally — the tier consumes the shared classifier directly. The
remaining holds are the operator's safety-contract review and the disabled-by-default flip.
Disabled by default
The tier exists only while the new
babysit_autopilot_merge_tieruserConfig (boolean, defaultoff) is enabled; the skill wires the
--autopilot-merge-tierflags only then. Enabling theflag, and any later gate-off flip, is a separate, announced operator step. Absent the flag the merge
gate is byte-for-byte its prior self, so
worker/autopilot's existing gate-proven merges areunchanged. The umbrella flag is fail-closed: it refuses (exit 3) unless
--lane-logins,--approver-bot-logins, and--block-labelsare all supplied.safety.md rationale
reference/safety.md's "Never do automatically → merge" contract is updated deliberately to codifythe tier and its criteria as an explicit, config-gated carve-out: the fleet may generate its own
approving review and merge only under the enumerated criteria, fail-closed, with a genuine bot
review and an untouched ruleset. This is the safety-posture change the operator signs off by
merging.
Test plan
python -m unittest discover -s tests— 287 tests pass, including newtest_babysit_merge.py(each criterion with a passing and a fall-back fixture, the decision-default veto's
pass / unratified / fetch-error cases, plus the tier-absent no-network invariant) and guard /
skill-contract additions for the fail-closed CLI and tier prose.
engine.test.sh— unittest suite + ruff clean + guarded-wrapper checks (including the newwrapper-level tier fail-closed) all pass.
babysit-readiness-gate.test.sh— unchanged, green (Python-free degrade + convergence intact).Related
Closes #476