Skip to content

test(coil): auto-resume tests wait on receipts, never on scheduler turns - #143

Merged
radroid merged 4 commits into
mainfrom
coil/134-autoresume-receipts
Sep 13, 2026
Merged

radroid merged 4 commits into
mainfrom
coil/134-autoresume-receipts

Conversation

@radroid

@radroid radroid commented Sep 12, 2026

Copy link
Copy Markdown
Owner

apps/server/src/coil/autoResume/Reactor.test.ts went red in CI on pending must be cleared: expected 1 to equal +0. Nothing was wrong with the product: advancePastResume was eight TestClock.adjust steps each chased by a fixed ten-pump spin — a budget of scheduler turns rather than a signal — and the store persists through writeFileStringAtomically, real filesystem I/O that completes on the Node event loop and not on TestClock. Under load the assertion simply looked before the cancellation had landed.

The auto-resume reactor now announces its milestones exactly the way the loop supervisor does (coil/loop/receipts.ts): an optional service resolved with Effect.serviceOption, a no-op Effect.void wherever nobody provides it — which is every production graph, since coil/index.ts does not mention the module — and a dropping PubSub so a test that stops draining can never stall the wake fiber. The variants are resume.scheduled, resume.skipped, resume.cancelled, resume.fired, and a tick.completed that closes every wake pass including the one where nothing is due, so "advance one poll and let the reactor finish" is one exact await. tick.completed is the only hot-path receipt and is assembled only when enabled, so an install with no subscriber allocates nothing per poll.

Every wait in the tests is now an await on a receipt, including the two negative ones: a disabled thread is proven by its resume.skipped rather than by a spin, and "it must not fire yet" is N demonstrably completed passes rather than N clock steps and a hope. The clock still moves in poll-sized steps — advancing time is the scenario, polling for the result was the bug — and settleQuiet / settleUntil / advancePastResume are gone.

Verification, all from the worktree: vp test run src/coil/autoResume (117 tests) ran clean 6 times after the change, Reactor.test.ts alone 5 times; the package typecheck (tsgo --noEmit) and vp lint on the three touched files are clean. A control run on the pristine base reproduced the flake class — 1 of 4 whole-directory runs went red on Reactor.test.ts.

Known and deliberately out of scope: the three autoResume/replay/*.test.ts suites still wait on settleQuiet / advanceSteps turn budgets from replay/reactorHarness.ts, and one of them (subagentFanout.test.ts) flaked once under parallel load during this work. Same class, same fix, now mechanical — worth its own change.

Closes #134

Claude Opus 5 via a Claude Code subagent did the work.

🤖 Generated with Claude Code

`Reactor.test.ts` went red in CI on `pending must be cleared: expected 1 to
equal +0`. Nothing was wrong with the product: `advancePastResume` was eight
clock steps each chased by a fixed ten-pump spin, a budget of scheduler turns
rather than a signal, and the store persists through real filesystem I/O that
completes on the Node event loop and not on TestClock — so on a loaded runner
the assertion simply looked before the cancellation landed.

The auto-resume reactor now announces its milestones the way the loop
supervisor does: an optional receipt service resolved with
`Effect.serviceOption`, a no-op `Effect.void` wherever nobody provides it
(which is every production graph), and a dropping PubSub so a test that stops
draining can never stall the wake fiber. `tick.completed` closes every wake
pass, including the one where nothing is due, so "advance one poll and let the
reactor finish" is one exact await.

Every wait in the tests is now an await on a receipt. The clock still moves in
poll-sized steps — advancing time is the scenario, polling for the result was
the bug — and the pump helpers are gone.

Closes #134

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

@radroid radroid left a comment

Copy link
Copy Markdown
Owner Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Reviewed the diff, receipts.ts against coil/loop/receipts.ts, and ran the tests. One real bug found (would use request-changes, but GitHub blocks that on your own PR); everything else checks out.

1. Possible hang: tick.completed is skipped if any fireOne fails, unlike the loop reactor it claims to mirror. Reactor.ts:329 — Effect.forEach(due, (p) => fireOne(p, nowMs), { discard: true }) has no per-item failure isolation, so if one fireOne fails, processDue fails before reaching the tick.completed emit at Reactor.ts:334. fireOne (Reactor.ts:213) calls snapshotQuery.getSnapshot(), which has a typed ProjectionRepositoryError channel — a real, expected failure, not just a defect — and only the dispatchResume call is wrapped in catchCause; nothing else in fireOne is. Compare coil/loop/Reactor.ts:700-707, where evaluateOne is deliberately wrapped per-thread in catchCause ("a single bad record must not retire it") specifically so the tick's end-of-pass receipt is unconditional. AutoResumeReactorReceipts has no such isolation.

Concrete failure: a snapshot read fails for one due thread mid-tick → processDue fails → the outer catchCause+forever wake loop logs and retries, so production is fine → but tick.completed was never published for that tick, so any test awaiting it via advanceOneTick's while (true) { PubSub.take(log) } (Reactor.test.ts) blocks forever, bounded only by vitest's default timeout instead of the old bounded-pump Effect.die with a clear message. Not hit by the current 8 scenarios (the stub getSnapshot never fails), but it's a latent hang the old code couldn't produce, and belongs in this PR since it's exactly the "same class" this PR is fixing.

Fix: wrap the fireOne call inside Effect.forEach with its own Effect.catchCause (log + continue), matching evaluateOne's pattern, so one bad record can't suppress the tick receipt.

Everything else held up:

  • All turn-budget waits (settleQuiet, advancePastResume, fixed pump loops) are gone; the two remaining TestClock.adjust calls each pair with an exact receipt await.
  • Negative assertions are now N-completed-ticks-observed, not spin-and-hope.
  • No production layer provides AutoResumeReactorReceipts (coil/index.ts never references it); receiptEmitter is a constant Effect.void/enabled:false when absent, and tick.completed is gated behind receipts.enabled.
  • receipts.ts matches loop/receipts.ts's pattern (buffer 4096, PubSub.dropping, scope/Layer.effect shape) with no drift.
  • resume.fired is emitted after the dispatch attempt (post-catchCause), so a failed dispatch still announces — correct as described.
  • Reactor.ts diff is additive-only around existing guards; no reordered/changed decisions. The 5 removed store bindings are all genuinely unused after the receipt-based rewrite.
  • Tests: Reactor.test.ts 10/10 passing x3 runs; src/coil/autoResume full dir 117/117 passing once. No wall-clock timeouts introduced.
  • Fork-owned files only, conventional commit title, PR body ends with the model/harness line, replay/*.test.ts follow-up explicitly deferred.

Claude Sonnet 5 via Claude Code subagent.

`processDue` ran every due arm through a bare `Effect.forEach`, so the first
`fireOne` that failed took the whole pass down with it. `fireOne` reads a fresh
snapshot per arm and `getSnapshot` has a typed `ProjectionRepositoryError`
channel, so that is an expected failure rather than a defect — and it cost two
things: the rest of the batch never ran, and the pass never reached its
`tick.completed`. Production recovered on the next pass, but anything awaiting
that receipt was left waiting for one that would never come.

Each arm is now wrapped the way `coil/loop/Reactor.ts` wraps `evaluateOne`: log
the cause, continue the batch. The failed arm itself is untouched — a failed fire
reserves nothing and clears nothing — so the next pass tries it again.

The new test pins both halves. Without the fix it fails on the arm: the first
thread's failure starves the batch until the wake loop's own retry re-runs the
whole pass, by which point the arm has been consumed out of order.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
@radroid

radroid commented Sep 12, 2026

Copy link
Copy Markdown
Owner Author

Fixed in 89d522523d.

processDue now wraps each arm in its own Effect.catchCause — coil auto-resume: fire failed with the threadId and Cause.pretty, then on to the next arm — exactly the way coil/loop/Reactor.ts wraps evaluateOne. tick.completed stays at the end of the pass and is now unconditional. The failed arm is deliberately left alone: a failed fireOne reserves nothing and clears nothing, so the next pass retries it.

New test: one failed fire neither aborts the batch nor suppresses the completed pass. Two threads due at the same moment, the first one's snapshot read fails once (the harness's getSnapshot stub grew a one-shot failure flag), and it asserts thread-2 fires in that same pass, thread-1 stays armed with zero attempts burned, and thread-1 then resumes on the following pass — ["thread-2", "thread-1"] in dispatch order.

Worth recording: reverted against this test the failure is not the hang you predicted, it is an assertion. The wake loop's catchCause+forever skips its own sleep on failure, so it immediately re-runs the whole pass; thread-1 fires on that retry ahead of thread-2 and the arm is gone by the time the test looks — expected [] to deeply equal [ 'thread-1' ]. The hang needs a persistent failure, which would have made a poor control. Either way the batch was being starved, so the fix is the same one.

Verification: Reactor.test.ts 11/11 x3, src/coil/autoResume 118/118, tsgo --noEmit on apps/server clean, vp lint on both touched files clean.

The rejected-window fixture keeps the sync's runtime.warning shape and this
branch's eventId parameter.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
@coderabbitai

coderabbitai Bot commented Sep 13, 2026 •

Copy link
Copy Markdown

Warning

Review limit reached

Next included review available in 23 minutes.

Check out review usage here.

View limit details

Limit details: You’ve used the included review currently available.

You've used all free OSS reviews for now. Wait for the free limit to reset to keep reviewing this public repository.

Learn how review limits work.

Review configuration:

⚙️ Run configuration

Configuration used: Path: .coderabbit.yaml

Review profile: CHILL

Plan: Advanced

Run ID: d9658fa2-9ecc-43de-b123-52a9e806327d

📥 Commits

Reviewing files that changed from the base of the PR and between 94d55bf and 5d3360c.

📒 Files selected for processing (4)
  • apps/server/src/coil/autoResume/Reactor.test.ts
  • apps/server/src/coil/autoResume/Reactor.ts
  • apps/server/src/coil/autoResume/receipts.ts
  • apps/server/src/coil/autoResume/replay/incidentDoomedTurn.test.ts

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

…ot a turn count

CI run 34734274701 failed its CONTROL case with `expected [Array(1)] to
include 'Auto-resume cancelled: thread-advanced.'`: ten quiet spins after
the clock step were not enough for the cancellation to land on a loaded
runner. Both cases now advance the clock and then `settleUntil` the activity
they assert on has been dispatched, which is the harness's own rule for
anything that must appear.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
@radroid
radroid merged commit f0464dc into main Sep 13, 2026
2 of 3 checks passed
@radroid
radroid deleted the coil/134-autoresume-receipts branch September 13, 2026 04:08
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

test(coil): auto-resume reactor tests wait on scheduler turns, and flaked once in CI

1 participant