Skip to content

fix(crews): J5 stream daemons start from the event store's high-water mark - #352

Merged
bryantderosier merged 3 commits into
j5/mainfrom
j5/j5-stream-high-water
Sep 29, 2026
Merged

bryantderosier merged 3 commits into
j5/mainfrom
j5/j5-stream-high-water

Conversation

@bryantderosier

Copy link
Copy Markdown
Collaborator

Problem

Three J5 background workers took their event-stream start point from SELECT MAX(sequence) FROM orchestration_v2_events. Nothing has written that table since the upstream sync, because V2 events now live in orchestration_events, so the query returns 0 and every boot replays the whole event history (#349). On j5/main that already lets the Captain archive cascade replay an old Captain archive on restart and retire live Crews while their Captain stays live. #315 would widen it to replayed unarchive, settle, and unsettle events, so this has to land before #315.

What I changed

  • New J5 helper apps/server/src/j5/a2a/eventStoreHighWater.ts → readEventStoreHighWater: reads EventSinkV2.latestSequence(). On failure it logs and retries with capped backoff (250 ms up to 30 s), and never falls back to 0.
  • CrewSeatFinishNotifier.ts → runDaemon and CrewLaunchReporter.ts → runDaemon: start from that helper. The raw queries are gone, and so is the launch reporter's unused SqlClient.
  • SilenceDetector.ts → initializeCursor: seeds its durable cursor from latestSequence(). A failed read fails the seed, and the lifecycle daemon already retries that step.
  • runtimeLayer.ts: the J5 runtime layers get OrchestrationV2EventSinkLayerLive. It's the same layer object the orchestration runtime uses, so there's one sink instance.
  • A separate commit fixes a misleading comment on fix(crews): Crew notices report only what the platform measured #292's daemon-mode alert test (Jackson's optional nit there; fix(crews): Crew notices report only what the platform measured #292 had already merged). That test is an end-to-end smoke check, and the wake-count test is the regression test.

Why this shape

  • EventSinkV2 rather than the lower-level store: ThreadManagementService.streamStoredEventsFrom streams from the sink, so the start point and the stream share one sequence space. It's also what QueuedRunWatchdog already uses.
  • Retry instead of starting at 0: starting at 0 is the bug. Giving up would stop Crew reactions until the next restart.

Invariants

  • No J5 daemon opens its stream, or runs its boot reconcile, before it has a real start point.
  • The start point and the stream read the same service.

Surfaces

Surface Decision
Entry points Unaffected.
Clients Unaffected.
Providers Unaffected.
Contracts Unaffected.
Reverse states Not applicable.
Connection modes Unaffected.
Upstream files / FORK.md None; only J5-owned files. FORK.md's "resumes from its high-water mark" is accurate again.
Docs Unaffected.

Out of scope

  • An install whose silence cursor was already seeded at 0 by the bug has already caught up on that one replay; this only changes how new cursors are seeded.

Upgrade and data

None. The next boot after this lands starts each stream at the current high-water mark.

Verification

  • apps/server, apps/web, and packages/contracts typecheck (exit 0).
  • vp test run apps/server/src/j5: 1,061 passed, 1 skipped. runtimeLayer.test.ts: 47 passed.
  • New tests, all signal-based with no sleeps:
    • The finish notifier's and the launch reporter's streams open after the mocked latest sequence.
    • A failed first read opens no stream until the retry succeeds.
    • SilenceDetector seeds its cursor from latestSequence, including after a failed read.
    • With the three source files reverted, all five fail.
  • Lint and format are clean.

Review focus

  • The retry-until-success choice in readEventStoreHighWater, and whether a daemon should ever give up.

Closes #349

Claude Opus 5.5 via Claude Code in J5 Code

🤖 Generated with Claude Code

bryantderosier and others added 3 commits September 28, 2026 13:31
… mark

The Crew seat finish notifier, the Crew launch reporter, and the silence
detector took their stream start point from MAX(sequence) of
orchestration_v2_events, a table nothing writes since the upstream sync.
The query returned 0, so every boot replayed the whole event history, and
the Captain archive cascade riding the notifier's stream retired live Crews
by replaying an old Captain archive.

All three now read EventSinkV2.latestSequence(), the same store
streamStoredEventsFrom reads. The Crew daemons go through
readEventStoreHighWater, which logs and retries a failed read with capped
backoff instead of falling back to 0. The silence detector's failed first
cursor seed fails its init, which its lifecycle daemon already retries with
backoff. The production J5 A2A layers get the orchestration runtime's own
OrchestrationV2EventSinkLayerLive, so Effect memoizes one sink instance.

Closes #349

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…ression test

The comment claimed the daemon-mode alert test waits until the test times
out when the alert does not wake the worker. The worker's wake queue can
hold a spare wake, so it passes without the fix; the wake-count test is the
regression test.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Conflicts in CrewLaunchReporter.ts and its test with #313: kept both imports, and #313's two new reporter tests after this branch's fixture.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
@bryantderosier bryantderosier added vouch:trusted PR author is trusted by repo permissions or the VOUCHED list. jackson-direct Taken by Jackson + Astra outside the fleet methodology; lanes never staff these labels Sep 28, 2026
@bryantderosier bryantderosier self-assigned this Sep 28, 2026
@coderabbitai

coderabbitai Bot commented Sep 28, 2026

Copy link
Copy Markdown

Important

  • 🔍 Trigger review

This repository does not receive automatic reviews because it has fewer than 10 stars.

⚙️ Run configuration

Configuration used: Repository: Jacksondr5/j5code/.coderabbit.yaml

Review profile: CHILL

Plan: Advanced

Run ID: 97217cfb-88ea-46f5-a536-1e61ab1f74eb


Comment @coderabbitai help to get the list of available commands.

@Jacksondr5 Jacksondr5 left a comment

Copy link
Copy Markdown
Owner

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Posted by an AI agent on Jackson's behalf.

Approved. This is the right fix for #349. It reads the start point from the event store service itself, so the start point and the stream share one sequence space without repeating a table name. The wiring stays in J5-owned runtimeLayer.ts. 41 tests pass across the five affected files, and CI is green.

On your review-focus question: retrying forever with capped backoff is fine to keep. It's small, and it never falls back to 0, which is the actual bug. By the "repair beats edge-case machinery" principle, failing loudly would also have been acceptable, since a store that can't answer its latest sequence is broken anyway. Not a blocker.

This supersedes #353, which duplicated this fix and is now closed. Merge this before #315.

@bryantderosier
bryantderosier merged commit a772b1a into j5/main Sep 29, 2026
35 checks passed
@bryantderosier
bryantderosier deleted the j5/j5-stream-high-water branch September 29, 2026 13:32
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

jackson-direct Taken by Jackson + Astra outside the fleet methodology; lanes never staff these size:M 30-99 effective changed lines (test files excluded in mixed PRs). vouch:trusted PR author is trusted by repo permissions or the VOUCHED list.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

fix(crews): J5 stream daemons replay all history on every boot (high-water reads an empty table)

2 participants