docs(plan): record the Phase 4 acceptance run and hold the tag at DOING - #362
Conversation
The acceptance test ran end to end for the first time and found two defects that every prior green signal had hidden. All four lane artifacts plus the three-clean-cycle auto-resolve are demonstrated on real lane output, and the runbook's one empirically unverified choice — the bad-token value — is now confirmed to produce class=auth rather than the non-escalating other. The tag stays at DOING rather than advancing: SC4 is unmet, the write path is only fixed on the #359 branch, and #238's lane routing (role label plus the routed-advisory marker) was never implemented. Deliverables are enumerated and individually dispositioned per Phase 3's ENUMERATION FIX, so closing the phase cannot strand an unbuilt one. The canary deferral gains a dated re-evaluation trigger and a strengthened rationale: a workflow-validation skip is green, reviews nothing, and emits no class token at all, so an annotation-based aggregator cannot see it. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
markdownlint reads a line starting with `#238` as a heading (MD018). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
The aggregator's lookup queries open issues only and documents the design as supersede-not-reopen, which contradicts #238's third acceptance criterion rather than merely leaving it unexercised. Which side is wrong is a decision, so it is recorded as blocking closure instead of patched. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…ling Across all 146 runs of this workflow exactly one ever failed, and on main's code a cycle reaching action != none must fail at the upload — so every other run necessarily reported action=none. That argument is airtight where the earlier "twenty-five consecutive runs" phrasing was only a sample. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…ter grammar Adds what the acceptance run established after the first ledger pass: - the branch-vs-main qualifier, stated as a qualifier rather than a footnote, since reading branch evidence as product evidence defeats the test's purpose - the 92-second detail: the hysteresis is counter arithmetic, so auto-close proves recovery was observed three times, never that it held - the third defect (a validation skip emits no class token at all) and the two lesser findings beside it, tracked in #363 - lane routing tracked in #364; neither needs-human nor routed-advisory exists anywhere in the repository, so it is unbuilt rather than misconfigured - the close-condition conflict and the unexercised multi-repo shape, both recorded on #238 for adjudication rather than resolved here Also corrects PLAN's own documented deliverable line against the shipped emitter: cycle carries a third value `incident` and there is a `coverage=` field. The omission mattered — `incident` is the value that says the watchdog fired, so the documented enum was missing the one outcome that matters. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
#359 merged as 058ed1a and the full open-to-close incident cycle then ran on main: incident #365 opened by run 31095551306, and closed by 31096244305 after three covering clean cycles. The first pass had only ever proven the fix branch, so the acceptance evidence now covers the shipped product rather than code that never ran in production. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…ath fix Review of the fix found it pinned only its own input, leaving three further ways to break the same round trip without any run going red. All three are now asserted, and the ledger says so rather than leaving the fix reading as complete on its own. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
The never-executed argument was self-contradicted by the very re-demo runs the ledger cites: they report action != none AND succeed. The corpus is now pinned by CODE STATE — every main run before 058ed1a — so the argument holds as the schedule keeps firing, and the run count is demoted to a dated observation rather than load-bearing. Deliverable-line excerpts are reordered to match the emitter, which puts read-errors before cycle and coverage. The sixth bullet's two downstream comment actions were un-dispositioned. Both are now performed and recorded, each carrying the half that is easy to gloss: claude-code-plugins#1327's signature classifies as non-escalating other, so detection is proven while no incident opens; provisioning#215's class=runner is escalating in the taxonomy but emitted by nothing in production, so that silence would still be silent. Neither is comment-closed — both carry needs-human. Also drops the "not ours to change" overstatement here, matching the wording already corrected in the workflow and the pull request bodies. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
The re-demo runs it points at are recorded above the block, not below. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…ings
The blocking one was an enumeration hole I had not seen: the canary-property
bullet names FOUR coverage targets and the ledger judged one. Its headline
("asserts on REAL lane output") IS satisfied, which is precisely why a silent
three-quarters made the whole bullet read as covered. Each target now carries
its own verdict — the credential-death target genuinely covered, the SDK
instant-fail detected but non-escalating by design, and the two runner targets
not exercised because nothing in production emits that class at all.
Remaining findings:
- the pre-fix run count is now scoped to the pre-fix corpus and marked as a
dated observation (143 main runs, verified independently), so no sentence's
validity rests on a number the schedule keeps changing
- the opener said "two defects" over a section enumerating three, and its
unbuilt-deliverable count is replaced by the enumeration itself so the two
cannot drift apart again
- the label-absence claim is scoped to workflows and scripts; as written it
would have been falsified by this very diff, which puts both strings in prose
- REMAINING is now a list, and states plainly that this phase cannot be closed
out autonomously at all: everything left is either code or a human's
judgement, including the closure half of the sixth bullet
- the marker-comment evidence points at the run log and explains why the pull
request no longer shows it, so its absence does not read as fabrication
- "roughly 92 seconds" matched no measurable boundary; "under two minutes" does
Left deliberately untouched: the pre-existing verbatim quote of run
30571900637, which predates the coverage field and must not gain one, and the
corrected emitter grammar line.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
|
Bugbot is not enabled for your account, so this pull request was not reviewed. Enable Bugbot in the Cursor dashboard to get automatic reviews on future PRs. |
|
Claude finished @kyle-sexton's task in 2m 58s —— View job Review complete
SummaryThis is a docs-only PR ( Since this PR's central claims are checkable facts about the shipped code, I cross-verified the ledger's assertions against the actual sources rather than taking them at face value:
No correctness or documentation-quality issues found — the ledger's factual claims about the shipped code check out, the tag is correctly held at Note: this PR is already merged; this review is retrospective and does not block anything. |
|
Claude has reviewed this PR 1 time. The lane skips further automatic reviews after 5; deleting this comment resets the count. |
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: 893d16ec09
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
| NOT BLOCKING, recorded so it is not mistaken for a gate: org secret | ||
| `CLAUDE_CODE_OAUTH_TOKEN` visibility is still `all`, NOT the | ||
| selected-repositories flip. That flip is operator-only UI work and remains | ||
| PENDING. It blocked nothing here — under `all` the sandbox reads the org | ||
| credential and the selected-repositories list is legitimately empty, which is | ||
| exactly the runbook's Step 0 disposition. It is a hardening item, not an |
There was a problem hiding this comment.
Remove the pending secret-visibility flip
If an operator follows this purported hardening item, they will violate the security contract in CLAUDE.md:39-43, which explicitly defines all-repository visibility as deliberate and says never to “correct” this secret to selected repositories. Re-scoping it can omit current or future consumers, causing their Claude lanes to run without the credential and potentially produce the advisory green/no-review failure this phase is intended to detect; keep visibility: all as the required final state rather than recording a flip as pending.
AGENTS.md reference: AGENTS.md:L9-L11
Useful? React with 👍 / 👎.
Records the Phase 4 acceptance run in
docs/topics/claude-review-lanes/PLAN.md. Draft, because the ledger it writes asserts that Phase 4 is not closeable yet, and that claim should be checked before it lands.Why the tag is
[DOING], not[DONE]The acceptance test ran end to end for the first time. It did what an acceptance test is for — it found three defects and two spec conflicts that every prior green signal had hidden — so SC4 ("#228/#238 closed with pointers") is unmet and both issues stay open. They also carry
needs-human, which bars autonomous closure independently of the evidence.What is now done: the write-path defect is fixed and merged (#359 →
058ed1a), and the incident lifecycle is re-demonstrated onmain. What remains is tracked in #363, #364, and two adjudication items recorded on #238.Phase 3's ENUMERATION FIX is the binding precedent here: a phase's close-out must enumerate its DELIVERABLES, not just its checks. The ledger therefore lists the unbuilt ones and dispositions each, so advancing the tag later cannot strand one.
What acceptance demonstrated
All four lane artifacts plus auto-resolve, on real lane output:
review / reviewconcludedsuccessFailure class: authclass=authannotationapi_error_status: 401in the check-run annotationaction=update), 31083359365 (action=close) — eachread-errors=0 cycle=clean coverage=completeThe runbook's one empirically unverified choice — the bad-token VALUE — is now confirmed: the SDK reached the API and was rejected 401, yielding
class=authrather than degrading to the non-escalatingother.Re-demonstrated on
main, because the first pass only proved a branchEvery incident write in the table above ran on the unmerged branch of #359 —
maincould not write an incident at all. #359 has since merged as058ed1a, and the whole cycle was re-run onmain:main)read-errors=0 cycle=incident coverage=complete action=openread-errors=0 cycle=clean coverage=complete action=updateread-errors=0 cycle=clean coverage=complete action=updateread-errors=0 cycle=clean coverage=complete action=closecleanCycles: 3, recovery commentSo the acceptance evidence now covers the shipped product. Teardown is clean: no repo-level override remains, zero incidents open, and the org secret's
updated_atnever moved from2026-08-05T13:32:31Zacross either pass.The three defects
1. The write path had never once executed — fixed and merged (#359). The poll renders the incident body to the dot-prefixed
.claude-lane-incident.md, andupload-artifactignores hidden files by default, so the upload collected nothing andif-no-files-found: errorfailed the step.That is provable rather than sampled, and the corpus is pinned by code state rather than a run count so it does not drift as the schedule keeps firing. While
maincarried the pre-fix code, any cycle reachingaction != nonemust have failed at this upload. Everymainrun before058ed1asucceeded except the dispatch that forced this incident — therefore every one of those reportedaction=none, and the write path was never exercised. Runs at or after058ed1asit outside that corpus by construction: the fourmainre-demo runs above reportaction != noneand succeed, which is the fix working, not a counterexample. (Dated observation, deliberately not load-bearing: 143 such pre-fixmainruns as of 2026-08-06.) Each pre-fix run emitted theread-errors=0 cycle=cleanline PLAN identifies as the antidote to the≤ 1ceiling check, while the write path could not fire.2. Lane routing was never implemented — tracked in #364. #238's Contract requires the incident issue to carry the human-gated role label plus a
kind=routed-advisoryescalation-marker comment. #361 and #365 carried neither. Neitherneeds-human(the role label both #228 and #238 themselves wear) norrouted-advisoryappears anywhere in this repository, on either branch — so this is unbuilt, not misconfigured. The label is applied inside the write-gate's byte-pinned region, so this PR does not patch it.The operational point: an
authincident inherently needs a human at the provider layer, and the issue does not wear the label that routes it to one.3. A validation skip is a silent no-review the aggregator cannot see — tracked in #363. When
claude-code-actionskips itself on workflow validation it exits 0, the check concludes green, nothing is reviewed, and noclass=token is emitted. The aggregator's whole detection mechanism is that token, so this failure mode is invisible to it by construction. That is the #228 harm class, uncovered. The same issue records two lesser findings: a review-count comment that counted a review which never happened, and marker copy falsely asserting a push does not re-trigger the lane.Two spec-vs-implementation conflicts, both recorded on #238 for adjudication
Neither is patched here — which side is wrong is a judgement call, and #238's Contract is ratified.
The reopen criterion. #238's third acceptance criterion says "a second incident reopens the same marker-selected issue (no duplicate)". The aggregator deliberately does not do that; its lookup queries
state: "open"only and says so:nextState's only actions areopen/update/close/none— noreopenexists — and a test pins the opposite behavior explicitly. So a second episode opens a new issue. The "no duplicate" half still holds (never more than one open incident, which is what SC3 ceilings), but "reopens the same issue" is contradicted by design, not merely unexercised.The close condition. #238's Contract says the incident closes on the "first window whose review runs include a success and no
authclass". This phase specifies — and the code implements — three consecutive clean cycles. The stricter shape ran, so nothing is broken, but two ratified authorities disagree and #238's wording is the stale one.The sixth bullet's two downstream actions, now performed
Phase 4's sixth bullet requires more than closing #228/#238 — it also requires comment-closing claude-code-plugins#1327 with root cause and a pointer, and commenting provisioning#215 as folded into the taxonomy. Both were un-dispositioned; both are done, and both stay open.
needs-human, which bars autonomous resolution. The comment states the half that is easy to gloss: that signature classifies asother, which is non-escalating by design, so detection is proven while no incident opens. Whether the instant-fail signature deserves its own escalating class is exactly the human decision left on it.class=runneris an escalating class in the taxonomy, and nothing in production emits it — the only occurrences are the aggregator's unit tests. So that substrate-silence would still be silent today. Unpark trigger: caller-side selector-failure emission shipping.The canary property's four targets, each with a verdict
The canary bullet names four coverage targets, and its headline ("asserts on REAL lane output") is satisfied — which is exactly why a silent three-quarters made the whole bullet read as covered. Review caught that; each now carries its own verdict in the ledger.
otherby design, so no incident opensclass=runneris escalating in the taxonomy but nothing in production emits itclass=runnerselector-failure marker (3a)Known coverage gaps, recorded rather than claimed
repositoriesSeen: 1).Also recorded
class=authunambiguous instead of confoundable with a wiring gap.class=token at all, so an annotation-based aggregator is structurally incapable of seeing it. A synthetic canary is the only proposed mechanism that would.visibility: allthe sandbox reads the org credential and the selected-repositories list is legitimately empty — the runbook's own Step 0 disposition. Recorded as a hardening item, not an acceptance gate. The secret was never edited or re-scoped:updated_atstayed2026-08-05T13:32:31Zthroughout, and the forced failure came from a repo-level override that was set and then deleted.Sequencing
This PR only records state; #359 is already merged, so the ledger's narrative is true on
mainas written. Do not advance the Phase 4 tag on the strength of this PR — the ledger's own "REMAINING TO CLOSE PHASE 4" list is what gates that, and it still names #363, #364, the two #238 adjudications, and the multi-repo gap.It also corrects PLAN's own documented deliverable-line format against the shipped emitter:
cyclecarries a third valueincident, and there is acoverage=field the documented line never had. That omission mattered —incidentis the value that says the watchdog fired, so the documented enum was missing the one outcome that matters.No linked issue — this PR records evidence and closes nothing. #228 and #238 remain open by design, for the reasons the ledger states.
Related
058ed1a.fe1b880.mainre-demonstration.🤖 Generated with Claude Code