Skip to content

feat(playbooks): the Captain runs a Crew's playbook and each seat receives its steps - #384

Merged
bryantderosier merged 13 commits into
j5/321-crew-playbook-planfrom
j5/322-crew-playbook-runs
Sep 30, 2026
Merged

bryantderosier merged 13 commits into
j5/321-crew-playbook-planfrom
j5/322-crew-playbook-runs

Conversation

@bryantderosier

@bryantderosier bryantderosier commented Sep 29, 2026 •

Copy link
Copy Markdown
Collaborator

Problem

With #321 a Crew records the playbook it follows and which seat owns which step, but nothing runs it: a Captain couldn't start the Crew's playbook, no seat ever got its step, and Fleet couldn't show where the run was (#322). This PR lets the Captain start the Crew's playbook, hands each step's live prompt to the seat that owns it exactly once per landing, cancels the run when the Crew is archived, and shows the current step and who holds it on Fleet.

This PR is stacked on #383 (#321), which is stacked on #380 (#319).

What I changed

  • packages/contracts/src/j5/playbook.ts → optional PlaybookRun.crewInstanceId (also on PlaybookProgress), PlaybookStepDelivery { state: "delivered" | "captain" | "pending", seat, threadId }, PlaybookStepResponse.delivery, and the crew_not_linkable and delivery_pending error codes. packages/contracts/src/j5.ts → optional FleetCrew.playbookRun.
  • apps/server/src/j5/a2a/migrations/029_CrewPlaybookRuns.ts → crew_instance_id on j5_playbook_run and a new j5_playbook_step_delivery table, one row per landing keyed (run_id, request_id).
  • apps/server/src/j5/playbooks/PlaybookStore.ts → start takes an optional crewInstanceId (part of the request JSON, so replaying a key with a different Crew is request_conflict). For a Crew-linked run, start, next, back, and reselect write the landing row in the same transaction as the move. New reads latestLanding, activeRunForCrew, and cancelForCrew.
  • New apps/server/src/j5/playbooks/PlaybookCrewRelay.ts:
    • start links the run under crews.serialize(crewInstanceId) and refuses crew_not_linkable when the Crew is missing, archived, commanded from another thread, or follows a different definition.
    • deliver resolves the owner from live members at the first attempt, persists the target seat and thread before dispatching, and sends the notice with a commandId and messageId derived from runId:requestId. A step with no owner, a seat that never launched, or an archived seat is the Captain's, with no notice.
    • mutate checks the run's owner, then drains every pending landing oldest first before moving. A landing that can't be handed off yet returns delivery_pending and the run doesn't move.
    • Nothing retries in the background. A hand-off that a crash or a transient failure left pending is finished by the Captain's next step call or a retry of the same one, which drains it before moving.
    • currentDelivery is read-only and backs playbook_current and Fleet.
  • New crewStepNotice.ts → the <j5_playbook_step> notice with the live prompt in <step_prompt>, escaped.
  • playbooks/mcp.ts → playbook_start takes crew_instance_id; the step tools and playbook_current return delivery. playbooks/instructions.ts → one bullet: branch on delivery.state; only captain means do it yourself; pending means don't start it yet.
  • ArchiveCrewService.ts → archiving a Crew cancels its active run, before markArchived and again on the already_archived retry, so the MCP tool, the Fleet button, and the Captain cascade all cancel. Stop leaves the run alone.
  • crewGateNotice.ts / CrewLaunchReporter.ts → the launch report names the playbook and how to start it with crew_instance_id.
  • FleetReadsHttp.ts → each live Crew's playbookRun (position, total, live step title, delivery state, seat). Web FleetPage.tsx and fleet.logic.ts → "Step N of M: <title> · ", "· Captain", or "· handing off to "; a run whose YAML can't be read or no longer has the step shows "Step · needs attention" through an optional issue on FleetCrew.playbookRun (position 0), so older clients still decode it. Fleet refetches on playbook changes. PlaybookRunsSection.tsx → a Crew-linked run shows "Crew · ", which opens and focuses that Crew's group, including a retired one.
  • Docs: docs/user/playbooks.md ("With a Crew"), docs/j5/product/features/playbooks.md, docs/j5/product/a2a/agent-tools.md.

Before (#321's head: the Crew row shows its playbook but not the step, and the run doesn't link to the Crew):

Fleet before

After (the Crew row shows "Step 3 of 4: Review · Captain", and the run card links to the Crew):

Fleet after

Why this shape

  • The Captain drives; the platform delivers. The Captain still calls the ordinary step tools and advances after the seat reports back. The platform never advances on its own, and the Captain gets no notice for its own steps, because the step tool response already carries the prompt.
  • Exactly once rests on two layers. The landing row is written in the move's transaction and keyed by the request that caused it, so a retried request can't create a second landing. The dispatch uses a deterministic commandId, so if the server crashes after dispatching but before marking the row, the retry is replayed by the orchestrator's command receipts and the seat still gets one message.
  • No landing is ever skipped. Every move drains the run's pending landings first and refuses with delivery_pending if one can't be handed off. I chose blocking over silently dropping a step. playbook_cancel never drains, so cancelling always works.
  • Pending is its own state. Only state: "captain" means the Captain does the step; seat: null alone never does. Guidance, playbook_current, and Fleet all branch on state, so a pending hand-off is never shown as the Captain's.
  • The owner is fixed at the first attempt. The target is persisted before dispatching and reused on retry, so a retry can't change the thread behind a commandId. Ownership changes apply to the next landing; back and reselect re-deliver under current ownership.
  • No background retry. The plan finished pending hand-offs with a sweep at server start. I dropped it in review (thread): after a restart, the Captain's next step call, or a retry of the same one, drains the pending landing before the run moves, and the persisted target plus the deterministic ids keep that one seat message.
  • Archive cancels; stop doesn't. Archive is the end of a Crew, and a cancelled run is terminal, so unarchive can't restart it. Stop is a pause.

Invariants

  • Each landing (start, next, back, reselect) produces at most one notice to at most one seat, including across retries and a crash between dispatch and marking.
  • A Crew-linked run can't move while an earlier landing is unresolved.
  • A thread that doesn't own the run is refused before any pending hand-off is sent.
  • playbook_current and Fleet never dispatch.
  • Runs without crew_instance_id behave exactly as before, with no delivery field.
  • Archiving a Crew, by any path, leaves no active run linked to it.

Surfaces

Surface Decision
Entry points (chat, Settings, command palette, keybinding) changed: the Captain's playbook_start and step tools, archive through MCP, the Fleet button, and the Captain cascade, and the Fleet page; no Settings, palette, or keybinding surface
Clients (web, desktop, mobile) changed: web and desktop Fleet and Playbook runs; mobile unaffected because it has no Fleet page, and the Captain thread's board already shows the run
Providers unaffected: the step notice is plain text through the orchestrator, the same for every adapter
Contracts (packages/contracts) changed, additive and optional, in the J5-owned src/j5.ts and src/j5/playbook.ts
Reverse states archive cancels and unarchive leaves it cancelled (start a new run with the same crew_instance_id); stop leaves the run; pending clears on the Captain's retry or next step call
Connection modes (local, remote, tunnel) unaffected: delivery is server-side; clients read existing J5 routes
Upstream files / FORK.md none: every changed file is J5-owned; the receipt integration test imports upstream orchestrator layers without editing them. FORK.md case 45 now says a Crew-linked start, next, back, or reselect dispatches a message into the owning seat's thread, which starts or queues a turn; a thread run's advancement still creates none
Docs changed: docs/user/playbooks.md, docs/j5/product/features/playbooks.md, docs/j5/product/a2a/agent-tools.md

Out of scope

Upgrade and data

Migration 029 adds a nullable column and a new table. Existing runs read back with no Crew link and behave as before. Older clients ignore the optional fields. Rolling back leaves a Crew-linked run behaving like a thread run. Numbering: j5/main's peering work took J5 migrations 23–27, so #383's is now 028 and this one is now 029. A dev database that ran an earlier build of this stack already recorded 23 and 24 as ours and would skip the peering migration with that number, so reset or re-copy such a database before testing. Real installs never ran the old number.

Verification

  • vp test run apps/server/src/j5/playbooks/PlaybookCrewRelay.test.ts apps/server/src/j5/playbooks/PlaybookCrewRelay.integration.test.ts apps/server/src/j5/playbooks/crewStepNotice.test.ts apps/server/src/j5/playbooks/PlaybookStore.test.ts apps/server/src/j5/playbooks/mcp.test.ts apps/server/src/j5/playbooks/PlaybookHttp.test.ts apps/server/src/j5/a2a/ArchiveCrewService.test.ts apps/server/src/j5/a2a/CrewCaptainArchiveCascade.test.ts apps/server/src/j5/a2a/CrewStopService.test.ts apps/server/src/j5/a2a/crewGateNotice.test.ts apps/server/src/j5/a2a/CrewLaunchReporter.test.ts apps/server/src/j5/a2a/FleetReadsHttp.test.ts apps/server/src/j5/a2a/Migrations.test.ts apps/server/src/j5/a2a/runtimeLayer.test.ts apps/web/src/j5/fleet/fleet.logic.test.ts apps/web/src/j5/crew/crewNotices.logic.test.ts apps/server/src/j5/a2a/ClientReadsHttp.test.ts apps/server/src/j5/a2a/ThreadHomesHttp.test.ts apps/server/src/j5/a2a/CrewLaunchService.test.ts: 19 files, 174 tests pass.
  • PlaybookCrewRelay.integration.test.ts runs through the real orchestrator and command receipt store: a failure injected after a real dispatch and before marking leaves the row pending; the Captain's next playbook_next drains it and replays the same commandId, and the seat thread has exactly one message with the deterministic messageId, including when the YAML prompt is edited in between.
  • Unit tests cover:
    • delivery on start, next, back, and reselect;
    • an unowned step, a seat that never launched, and an archived seat going to the Captain;
    • a retried next delivering once;
    • delivery_pending blocking the move and draining in landing order;
    • playbook_current reporting pending rather than captain;
    • another thread's move refused before any pending hand-off is sent;
    • archive (direct, already_archived retry, and Captain cascade) cancelling the run while stop doesn't;
    • the Fleet projection and owner label.
  • No test waits on a sleep; they wait on the orchestrator drain and the landing row.
  • vp exec tsc --noEmit -p . in packages/contracts, packages/client-runtime, apps/server, and apps/web: clean. vp lint on the touched files: clean.
  • Screenshots come from an isolated dev server with a sample playbook and a two-seat Crew, with the run moved to the Captain's Review step. Web only.

Review focus

  • PlaybookCrewRelay.mutate and deliver: the owner check before the drain, the target persisted before dispatch, and which failures resolve as captain versus stay pending. I most want the exactly-once argument challenged.
  • ArchiveCrewService.archive: cancelling before markArchived and on the already_archived retry, under the Crew lock.

Closes #322 (stacked on #383)

Claude Opus 5.5 via J5 Code (crew: planner, builder, and two reviewers on Claude and Codex)

🤖 Generated with Claude Code

…eives its steps

A Captain starts the playbook its Crew follows with playbook_start(...,
crew_instance_id). Every landing (start, next, back, reselect) writes one
delivery row in the move's transaction, and the new PlaybookCrewRelay
hands the step's live prompt to the live seat that owns it, once per
landing, under the Crew's lock. Unowned steps, seats never created, and
archived seats are the Captain's, with no notice. The step tools and
playbook_current report delivery { state, seat, threadId }; only
state "captain" means the Captain does it, and a pending hand-off blocks
the next move with delivery_pending until it resolves. A boot sweep
finishes hand-offs a crash or a transient failure left pending.

Archiving the Crew (MCP, Fleet, or the Captain cascade) cancels its run;
stopping it doesn't. The launch report names the playbook and how to
start it. Fleet shows each Crew's step and who holds it, links a
playbook run to its Crew, and refreshes on playbook changes.

Migration 024 adds the run's Crew link and the delivery table.

Closes #322

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
@bryantderosier bryantderosier added enhancement New feature or request size:XL 500-999 effective changed lines (test files excluded in mixed PRs). labels Sep 29, 2026
@bryantderosier bryantderosier self-assigned this Sep 29, 2026
@bryantderosier
bryantderosier added this pull request to stack #385 September 29, 2026 22:47
@coderabbitai

coderabbitai Bot commented Sep 29, 2026 •

Copy link
Copy Markdown

Important

  • 🔍 Trigger review

This repository does not receive automatic reviews because it has fewer than 10 stars.

⚙️ Run configuration

Configuration used: Repository: Jacksondr5/j5code/.coderabbit.yaml

Review profile: CHILL

Plan: Advanced

Run ID: c0488c06-a887-43f4-a5d4-31364b55d607

  • Autopilot · Keep fixing CodeRabbit findings and required CI, and resolving merge conflicts

Autopilot is currently an internal CodeRabbit preview.


Comment @coderabbitai help to get the list of available commands.

@github-actions github-actions Bot added the vouch:trusted PR author is trusted by repo permissions or the VOUCHED list. label Sep 29, 2026
…sh statement

The SQLite client caches prepared statements by SQL text. The upgrade test
read `SELECT * FROM j5_playbook_run` before and after migration 024 with the
same text, so the second read reused the pre-ALTER statement, and on Node
24.14 (CI's .nvmrc) that statement keeps its old column list without
crew_instance_id. The post-migration read now names its columns.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

@Jacksondr5 Jacksondr5 left a comment •

Copy link
Copy Markdown
Owner

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

[Review panel: Opus 5.5, Astra, Sonnet 5.5, Sol 6.1]

The panel agreed on three findings, which are inline. There's also one FORK.md item, which can't go inline because FORK.md isn't in this diff:

FORK.md case 45 is now inaccurate. It says: "advancement never creates turns or controls provider lifecycle." With this PR, a Crew-linked playbook_start, next, back, or reselect dispatches a message.dispatch into the owning seat's thread. That message starts a turn, or queues behind the seat's active one. Please rewrite that sentence in case 45 in this PR, so the case says what Crew-linked advancement does.

Astra, Sonnet 5.5 and Sol 6.1 live-tested the stack tip, and all six scenarios passed: persona steps, proposal and roster, per-seat delivery, the Fleet step line, archive cancelling the run, and a plain Crew. Opus 5.5 reviewed the code and ran focused tests. Product questions go to Jackson separately. Those are where the Crew–playbook link lives, how big the delivery ledger should be, and whether the Captain should be told when an archive cancels its run.

Comment thread apps/server/src/j5/playbooks/PlaybookCrewRelay.ts Outdated
Comment thread apps/server/src/j5/playbooks/PlaybookCrewRelay.ts Outdated
Comment thread docs/j5/product/features/playbooks.md Outdated
bryantderosier and others added 7 commits September 29, 2026 23:07
A pending step hand-off is now finished only by the Captain's next step
call or a retry of the same one: drain-before-move, delivery_pending with
"retry the same call", the persisted target, and the deterministic command
and message ids are unchanged. The relay no longer forks a reconcile at
boot, so the layer has one shape. The crash-window integration test now
recovers through the Captain's next move and still proves exactly one seat
message through the command receipt.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
When an active Crew run's live YAML can't be read, or no longer has the
recorded step, Fleet used to drop the run. It now shows the recorded step
id marked "needs attention", with the reason on hover. The contract gains an
optional `issue` on FleetCrew.playbookRun (position is then 0), so an
older client still decodes the read.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…allback

Rewrites the Definition, AC5, AC12, and the release scenario for Crew runs
that hand each step to the one seat that owns it, with the Captain doing
unowned steps, and separates Crew-run archive (cancels) from a single
agent's (resumable). The 2026-09-29 History line now only records and
links the decision.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…s thread

Case 45 said advancement never creates turns. A Crew-linked start, next,
back, or reselect now dispatches a message into the owning seat's thread,
which starts or queues a turn there; a thread run still creates none.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
@bryantderosier

Copy link
Copy Markdown
Collaborator Author

@Jacksondr5 on the FORK.md item from your review: case 45 is rewritten in cccf51b. It no longer says advancement never creates turns. It now says a Crew-linked playbook_start, next, back, or reselect dispatches a message into the owning seat's thread through ThreadManagementService.dispatch, which starts a turn or queues behind the seat's active one. All three inline threads are fixed and resolved too, and #383's latest is merged in (d7ecd34).

…bered to 29

j5/main's peering migrations take J5 ids 23-27 and #383 moves
CrewPlaybooks to 28, so CrewPlaybookRuns becomes 29. The upgrade test
now runs through 28 before 29 and keeps its named-columns read. The
playbooks definition keeps its rewrite and gains main's pre-approval
sentence.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
@vercel

vercel Bot commented Sep 30, 2026 •

Copy link
Copy Markdown

The latest updates on your projects. Learn more about Vercel for GitHub.

Project Deployment Actions Updated
j5-code Ready Ready Preview Sep 30, 2026 5:56pm UTC

Request Review

FORK.md case 45 keeps #314's suggestion-seam wording and this branch's
Crew-linked dispatch sentence.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
@Jacksondr5

Copy link
Copy Markdown
Owner

[Review panel lead, Opus 5.5] The FORK.md case 45 item from the panel review is verified in cccf51b. The case now says a thread run's advancement creates no turns, while a Crew-linked start, next, back or reselect dispatches into the owning seat's thread through ThreadManagementService.dispatch. The wording survives the #314 merge alongside case 49.

@Jacksondr5 Jacksondr5 left a comment

Copy link
Copy Markdown
Owner

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

[Review panel: Opus 5.5, Astra, Sonnet 5.5, Sol 6.1] One low docs note from round 2 is inline. All round-1 findings are verified fixed (replies on each thread).

Comment thread docs/j5/product/features/playbooks.md Outdated

@Jacksondr5 Jacksondr5 left a comment

Copy link
Copy Markdown
Owner

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

[Review panel: Opus 5.5, Astra, Sonnet 5.5, Sol 6.1]

Approving at dea99cb. Every round-1 finding is fixed, and the panel checked each one against the source:

  • the boot sweep is removed;
  • Fleet's "needs attention" header now shows for a Playbook file that can't be read;
  • the definition is rewritten;
  • FORK.md case 45 is corrected.

C4, the "unowned step" wording, is a non-blocking docs nit. The size of the delivery machinery is deferred to #389.

This stack merges with #387, which is waiting on Jackson's decisions on C1 and C2.

instructions.ts keeps #381's mention bullet and this branch's Captain
delivery bullet.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…ain's

The release scenario said a step naming no Role is the Captain's. What
decides it is ownership: a step no seat owns is the Captain's, whether or
not it names a persona.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
@bryantderosier
bryantderosier merged commit 08aee52 into j5/main Sep 30, 2026
37 checks passed
@bryantderosier
bryantderosier deleted the j5/322-crew-playbook-runs branch September 30, 2026 18:19

This branch was successfully deployed

1 active deployment
Preview — 24a2635a Deployed Sep 30, 2026 by vercel[bot]
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

enhancement New feature or request size:XL 500-999 effective changed lines (test files excluded in mixed PRs). vouch:trusted PR author is trusted by repo permissions or the VOUCHED list.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

The Captain runs a Crew's playbook and each seat receives its steps

2 participants