Skip to content

Harness keeps thread state true: turn lifecycle, wake turns, lost-work recovery, watched work, containment, durable log, flow control - #40

Merged
i2cjak merged 57 commits into
mainfrom
backplane/harness-final
Oct 3, 2026
Merged

i2cjak merged 57 commits into
mainfrom
backplane/harness-final

Conversation

@i2cjak

@i2cjak i2cjak commented Oct 3, 2026

Copy link
Copy Markdown
Owner

The harness, not the model, now keeps a thread's state true and its work going. Reviewed adversarially: two Sonnet 5.5 and two GPT-6.1-Sol reviewers, pair consensus, then an Opus 5.5 / Sol debate wrote the spec. Implemented by Opus and Sonnet agents, then checked by two more Sol review rounds. Their 20 findings were fixed with laws and tests.

Root cause of "complete while working" (reproduced from the live hub's log): Claude Code runs a queued task-notification as its own turn, and the hub took that turn's result as the end of the user's turn. Real CLI captures (test/fixtures/) settled the rule: each message gets a uuid, and the turn ends only on that command's command_lifecycle completed|cancelled. A self-started wake turn shows as running. background_tasks_changed is the open-work list.

Turns finish and state stays true

  • One shared busy check drives alerts, thread_wait, settle, scratch expiry, the island, delegated tasks and bot replies. Failed now outranks Monitoring.
  • Lost work is told and resumed after an agent exit, a restart or an install. Deferred restarts honour restart.continue.
  • Spawn generations: a late start, input or MCP call from an old turn is dropped. Kills are guarded.
  • Exit is detected without stdout EOF (Claude, Codex, Grok), and the final partial line is flushed.
  • work_run / work_watch / work_list / work_cancel / work_unwatch: the hub owns long jobs, the outcome survives restarts, and cancel checks the process identity first. Agents are told to end their turn instead of polling.
  • Agent containment: each agent runs in its own memory-limited systemd-run --user scope (agent.memory). An OOM kills only that agent, and the turn says so. Scopes are tagged with their hub's home, so a test hub never stops the live one's agents. --expand-environment=no keeps prompts literal.
  • Installs (dev.sh install, in-app updates) check for busy threads immediately before the restart; a forced update stays distinct.

Durability: one write per step. A failed append is cut back to its pre-write length, and commits are refused until that works. A torn tail is repaired. MCP replies go out after commit. Delegated results are delivered idempotently (the wake is written before TaskDelivered).

Instant and exact updates: send deadlines renew on progress. Every producer goes through flow control, and the bootstrap is chunked by bytes (iOS 64 MiB safe; a split frame says more:true). A dropped client's channel is drained, and each frame is encoded once. Heartbeats run every 15 s (all four clients). Live text survives reconnects as a replacement snapshot. Bot RPCs are idempotent. Assistant blocks get distinct ids. Far mirrors retry in-flight batches and carry work rows. Only the thread this client created gets selected.

Many projects: the sidebar index is built once per frame, and phone alerts compare only affected threads. Progress writes are coalesced. Fr.forget runs in one pass. Turn preparation (K.Files.read, skills) moved off the hub loop. The window lays out a streamed reply's settled prefix once.

Tests: scripts/check.sh (All terms check, ~200 new laws, none from main removed or weakened), scripts/test.sh 84/84, native build, APK + bridge.js. New e2e all ok: wake, restart_work, restart, durable, logfull, spawn_gen, exit, prep, watch, live, contain, plus retire, stall, emb, proj_thread, far, bots, watch_scripts.

Not done: the streamed-text layout cache is native-only (the web and phones show the same text; it only saves layout work). Cancelling Claude Code's own background tasks is left out because no CLI control request for it exists. No device runs yet: a TestFlight build follows the merge.

i2cjak added 30 commits October 2, 2026 11:45
…urns

Item 1 of the harness-continuity spec. src/core/cmds.bend is the pure
command ledger (laws cmds_*, wake_act_once): every Claude user line
carries a uuid minted when the hub takes the message (Hub.uuid,
Claude.user.u; kept by the server's rewrites), the server keeps a ledger
per Claude process (ARun) and judges each stdout line before the step
(Hub.agent_line.v): a result never ends a turn by itself, the end comes at
the command's command_lifecycle completed/cancelled with nothing else
outstanding, a task_notification (or a bare init) starts a visible wake
turn, Stop makes unanswered commands tombs, and a CLI older than 2.1.286
keeps the old rule (verdict 0 is exactly Hub.agent_line). Messages whose
images are read off the hub are written in the order they were taken.

Tests: test/claude_cmds_test.bend replays both recorded 2.1.286 streams
(test/fixtures/claude/) line by line plus the live repro and Stop cases;
bun test/tools/wake_e2e.ts runs the stand-in claude
(test/tools/claude_standin.ts) through wake, stale-first, stop, pair, bg
and noLC.
… mirrors

Item 2. A background task that finishes (task_updated, or gone from
Claude's background_tasks_changed list: Bg.sync, which never adds) keeps
its row as ended (M.Bg.end, doing \u0001ended:<status>) until its
task_notification hands the outcome to the agent, or the next turn end
settles it (Bg.settle); it still counts as work, and Subs shows it as
"reporting" / "finished, reporting" on all four clients. Rows are
written as the new WorkSet change (folded exactly as SettingSet
<kind>.<thread>; old logs keep decoding), which far.bend shares with a
thread's mirrors (Fr.shared/Fr.change); Fr.forget drops a link's work
settings. Prompt: <background_work> says Backplane wakes the agent.

Laws bg_sync_*, bg_updated_ends_not_drops, bg_notification_drops,
bg_empty_snapshot_keeps_pending, bg_turn_end_settles_ended,
workset_fold_same_as_setting, store_workset_encoded, far_workset_*,
far_forget_drops_work, prompt_background_work. Tests: claude_cmds_test
(rows after every fixture line), subagents_test, far_test; wake_e2e bg
and far_e2e (a yielded task reaches the mirror).
…tasks and bot replies

Item 3. model.bend's Work.busy (unfinished child tasks, background rows,
ended ones included, and watched work) and Work.ready (not running,
nothing queued, nothing out) are the one check: alerts
(Notice.decide.w: a finished turn with work out tells nothing, a failure
or a question still alerts), thread_wait (Mcp.done.w, with "waiting" and
"work"), settling, scratch expiry and worktree removal on the tick (the
children counted once, Work.kids.map), the phones' island (Note.island.w,
"waiting on N"), delegated tasks (Task.finish.w judges the step's last
turn change against the state after it: a child's task waits for its own
work and queued follow-ups; failures end at once) and bot replies (the
answer to a person elsewhere is captured as BotReply by the step that
leaves the thread ready, and sent once via the Replied message instead of
bots.answered re-reading the log later). A failed turn shows Failed before
Monitoring (Phase.of.w, Status.wait.f).

Laws work_*, notice_busy_*, notice_decide_w_idle_same, wait_*,
settle_waits_for_work, settle_without_work, scratch_busy_kept,
worktree_busy_kept, island_busy_shows, task_*, bot_reply_*,
phase_failed_before_monitoring, status_wait_f_keeps_failed. Tests:
notice_test, subagents_test (child, grandchild, root; queued follow-up),
scratch_test; wake_e2e (no notify until the wake), bots_e2e (one reply).
…epair a torn tail

Item 4. The server writes a step's lines in one write (ST.Store.text) and
folds, announces and runs the step's effects only once it is written; a
failed write folds nothing and runs nothing, is logged and counted
(hub.write_fail) and shown to every client (update error "log write
failed"), and the hub keeps serving. An MCP tool call is answered after
its step is stored, and told when it could not be. At start-up a log
whose last byte is not a newline gets one (ST.Store.tail_fix), so a torn
fragment stays one skipped line. A delegated task's result reaches the
parent under a message id of the task's (Task.msg via Hub.inbox.id) and
TaskDelivered comes last; Hub.recover finishes torn deliveries
(Task.repair: a parent holding the message only gets TaskDelivered, and an
idle one with it newest starts its turn; a task whose child's turn ended
before its state was set is finished; otherwise the result is delivered).

Laws commit_text_joins, task_finish_delivered_last, task_repair_*,
store_torn_tail_resyncs, store_torn_tail_newline (proof of
inbox_never_interrupts updated for the message id argument). Tests:
test/store_torn_test.bend cuts the delivery step after each of its 449
bytes and recovers (one result, delivered, parent running every time);
bun test/tools/durable_e2e.ts SIGKILLs the hub as a child finishes and
appends a torn line.
…elease, install waits

Item 5. Background work whose owning process is gone is told to the agent
once (Bg.lost: "[Backplane] <why> ended your background work. Ownership
was lost; the outcome is unknown.", one line per task, "Check what
finished before rerunning.") and its rows are cleared in the same step;
with restart.continue off it is an Act "lost" and the rows stay, ended,
outcome unknown. Why: a restart (Hub.recover marks the thread,
resume.<thread>, and Hub.carry tells it, Rec.burst() = 4 threads started
at most; the rest go each minute and after each turn end through
Rec.release and the server's Release message; a marker that outlives a
second restart is only logged), the agent's exit (Hub.agent_gone; behind
the failed turn's queue when one ran), or its replacement for a new model
or mode (committed synchronously before the old process stops, so the new
process's rows are untouched). A held queue is never drained by a lost
note. A restart no longer fails a delegated child whose turn carries on
(Rec.turn.k with Carry.will); Task.repair judges children from the state
the hub started from. Updates wait while threads run or wait on work
(update.wait, released by the tick when they are done or after 30 min, or
by the new update.now RPC); Settings on all four clients says "Update
waits for N threads" with Update now; /hello reports "busy", and
scripts/dev.sh install refuses while it is above 0 unless --force.

Laws recover_lost_*, recover_no_work_quiet, lost_off_logs_and_keeps,
gone_idle_told, gone_running_queues, replaced_*, carry_keeps_child_task,
carry_off_fails_child_task, carry_child_result_once, rec_burst_bounded,
rec_release_bounded, rec_release_never_drains_queue,
lost_text_not_carry_text, update_wait_*. Tests: test/restart_test.bend
(ten threads, repeated restarts, continue off, a child with work out);
bun test/tools/restart_work_e2e.ts (kill by PID and restart, a child
carried across it, a replacement, continue off, /hello busy and the
dev.sh install refusal).
Item 6. Every agent process the server starts is a generation (Hub.gen:
when it started and the log's length; minted by the server as it handles
AgentStart, CodexTurn and GrokTurn). A Codex or Grok turn registers a
placeholder agent keyed spawn:<gen> (pid 0) at once; its TurnPid is
adopted only by that generation (B.Gen.accepts), any other is stopped, and
a failure to start comes as SpawnFailed, heard only from the current
generation. A Claude process keeps its generation in ARun; a message whose
images were read for a generation is written only to it and only if it
was not stopped (B.Gen.write). Every agent's MCP URL is
/mcp/<thread>.<gen>: a tools/call that changes something from a process
that is no longer the thread's is refused as a stale process, reading
tools still answer (B.Gen.refused). No pid of 0 or 1 is ever signalled
(Proc.kill through B.Proc.kill.ok; the voice helper too).

Laws kill_never_low_pid, spawn_gen_accepts_current, spawn_gen_drops_stale,
stale_spawn_fail_quiet, mcp_stale_gen_refused, mcp_stale_gen_reads,
prep_stale_dropped, gen_mcp_path. bun test/tools/spawn_gen_e2e.ts: codex
turns (one stopped while starting and sent again), grok that cannot start,
a replaced Claude process's MCP calls refused or answered.
Item 7. Each Claude process has a watcher fiber (Agent.watch) that asks
each second whether it exited, without reaping it (new effect
Proc.exited: waitid with WNOWAIT; the JS twin, which runs no agents,
always answers running). On exit it tells the hub (AgentExited, ignored
from a stale generation: B.Exit.current) and, if the reader has not seen
end of stream 2 s later because a child holds the output, shuts the
reading side of a dup of its socket (new effect Sock.shut_read) so the
reader ends; the reader reads a last line without a newline
(Agent.last, U.Lines.rest) and handles the exit once. The watcher ends
with its process. A running Claude turn with no line for 10 minutes and
no work out gets one Act "quiet" (B.Quiet.due, ARun.heard/quiet, the
server's Silence message each minute); nothing is killed.

Laws quiet_once, quiet_never_when_busy, quiet_never_idle,
exit_stale_gen_quiet, final_partial_line_flushed. bun
test/tools/exit_e2e.ts: exit with a child holding the output (fails in
about 3 s), a completed command then that exit, a last line without a
newline, a plain exit handled once, a replaced process's exit ignored.
…el, waiting contract

Item 8. New MCP tools work_run {command, note, timeoutSec}, work_watch
{kind: unit|pid|file, target, note, timeoutSec}, work_list, work_cancel
and work_unwatch (mcp.bend over hub.bend's Watch.*). A row (WorkSet kind
"watch") is written before anything is launched and the Await to launch
and watch it goes out with that step, so the tool's success follows the
write. The server runs a work_run as <home>/work/<id>/run.sh, as the
transient unit bp-<id> (systemd-run --user, so it outlives Backplane's
own service) or set apart with setsid without user systemd; its exit code
is written whole to code (code.tmp renamed), the outcome even across hub
downtime, and a started marker keeps a relaunch after a crash from
starting it twice. A watcher checks every 2 s: a unit by its
InvocationID, a pid by boot id and start time, a file by a token it must
contain; gone, reused or past its deadline reads "outcome unknown",
never success. Its end is told once (the row's removal): an idle thread
wakes, a running one gets it queued; a timeout leaves the row ended,
unknown. work_cancel shows cancelling until the end is confirmed
(cancelled); work_unwatch says the work may still run and stops the
watcher. Watched work counts as the thread's work (M.Work.busy), shows in
the Subagents panel with a Stop on all four clients (actions work-cancel,
task-cancel; RPCs work.cancel and task.cancel, a delegated task's parent
told once), and a restart watches open work again (Watch.rearm); a mirror
never watches. Prompt: start long work with work_run or register it with
work_watch, then end the turn; never poll or sleep-wait.

Laws watch_intent_before_launch, watch_end_once, watch_end_wakes_idle,
watch_end_queues_running, watch_timeout_unknown_not_success,
watch_pid_identity_unknown, watch_cancel_keeps_until_confirmed,
watch_cancel_confirmed, watch_unwatch_says_may_run,
recover_rearms_watches, mirror_never_awaits, prompt_background_work_b.
Tests: test/watch_test.bend; bun test/tools/watch_e2e.ts (exit 4 with the
output's end, a file's token, an outcome retained across a hub kill, a
pid gone, work_cancel, a Subagents Stop, a timeout).
Sub-items (c), (d) and (e) of item 9 land together (they share the
Client record and the join path):

- Every frame for a client goes through Wire.offer and is counted
  (core/outbox.bend's Cr); writers (and the window's Fwd) report what
  they put out (Drained, per connection generation). A live client more
  than 96 frames or 1 MB behind goes back to taking the log's changes in
  chunks from where it is (Cr.catch) instead of being dropped for a
  burst; one more than 768 frames or 8 MB behind is cut at once through
  a raw descriptor (Sock.raw/Fd.shut). Wire.all returns the clients.
- Pongs go through the hub (credit) at most once a second per
  connection; the reader scans pings, text frames and pongs in one pure
  pass (Conn.scan) and Ws.decode no longer measures the whole buffer
  per frame (a 20000-ping flood froze the hub 7 s).
- A join sends its first 64 lines with "more", then a chunk per
  JoinMore (one in flight, a couple of ms apart so the loop polls
  sockets), then each far mirror in chunks keyed by sequence, then a
  last frame with "boot": false. The hub never reads the lines' text: the
  writer puts each chunk together (WItems). Clients hold "@boot" while a
  log arrives in chunks and phones raise no alerts for it.
- The hub reads two inboxes: requests, joins and reports straight in;
  agents' lines, jobs and its own posts behind 16 tokens (Bulk.*), so
  Stop during a flood lands in ~100 ms.
- A broadcast to two or more sockets is encoded once (WBin); a single
  socket's frame is encoded by its writer; the window gets text.

Laws credit_*, boot_*, join_*, outbox_fifo, pong_rate_bounded,
wire_all_bin_eq, ui_chunked_join_equals_whole. test/hubs_test.bend: a
split join raises no alerts. stall_e2e.ts: slow join of a 24 MB log,
non-reading join, ping flood, resumed join, Stop during a flood, memory
over dropped joins; requests answered throughout (worst 100 ms).
A thread's streamed text is not in the log until its reply is posted,
so a client that reconnected mid-reply appended the rest to what it had
(abc + ghi) or lost what streamed while it was away.

- The core emits Out Live{thread, text} for streamed text (Delta.push);
  the server sends it as the delta it was and keeps it per thread
  (core/live.bend's Buf in botnet's Bn: chunks newest first, capped at
  256 KB with the cut noted). A step that posts an assistant message or
  ends the turn clears the buffer (Live.ends, after the commit).
- A join's last frame (one-frame joins, chunked joins' Boot.last, the
  window's WState join) is followed by {"t":"live","items":[{thread,
  text, trunc, gen}]}. Each clear bumps a buffer's gen; a new buffer
  starts at the hub's clock x 1000, so a restarted hub never numbers
  below one a client took.
- Clients (shared Ui: web, phones through Hubs, the window through the
  same recv after its WState): Ui.live.set replaces the live text with
  the snapshot ("…" first when cut), clears threads not in it, and
  ignores an item older than the gen it last took for that thread. Gens
  are compared as decimal strings (Nat.read makes laws crawl).

Laws live_snapshot_replaces, live_snapshot_idempotent,
live_snapshot_clears_absent, live_end_clears, live_turn_end_clears,
live_joined_exact, live_old_gen_quiet, live_trunc_marked.
test/live_test.bend; e2e test/tools/live_e2e.ts (stand-in claude mode
"stream"): reconnect mid-reply shows it exactly and whole, a reply that
ended while away shows no live text, a client back during a second
reply shows only that one.
Ui.see selected whichever ThreadCreated came next while want_new was
set, so another client's new thread (or a bot's, or a replayed log)
could take the selection.

- Hub: Reply.made(client, id, thread) answers thread.create, thread.fork
  and project add/new (through thread_create); Reply.ok is unchanged and
  a brought-back project's answer names no thread.
- Client (shared Ui, all four clients): New Thread, Fork and adding a
  project keep the request id in "@want.new" and where it went
  ("@want.new.at": machine|generation). Made.reply: only the answer to
  that request, on the same hub and connection, wants the named thread
  (Want.thread: selected at once if here, else when it arrives); a
  failure, a newer request or an older hub's plain answer selects
  nothing. Ui.see's creation arm is gone (the def stays); the want_new
  field is no longer set (no Ui type change, phone key unchanged).

Laws want_new_ignores_other_threads, made_selects_own,
made_selects_own_later, replay_cannot_select, made_stale_quiet,
made_fork_selects_own. test/picker_test.bend updated (the answer names
its thread; a plain answer and another's thread select nothing);
proj_thread_e2e.ts: answers name their thread, two clients at once,
a fork.
Side.ix (never kept) gathers open asks per thread, live child tasks per
parent and threads by project once a frame or phone screen; Side.phase.ix,
Subs.count.ix, Side.rows.ix, Active.rows.ix answer from it, threaded through
the window, web and phone screens (Active parts once per phone screen). Phones
compare Seen only for the threads a changes frame touches. Laws side_rows_ix_eq,
side_phase_ix_eq, subs_count_ix_eq, active_rows_ix_eq, settle_ix_eq,
alerts_ids_same, alerts_ids_raise. Bench test/native/side_ix_bench.bend (JS,
100 projects, 3000 threads, 5000 asks, 800 tasks): old 13 s a frame, ix 0.26 s.
Stage 2 (persisted index) not done: it bumps the phones kept key.
…ns the hub

A pushed batch is kept until the other end answers; a failure puts it back
ahead of what waited and holds the link for a 2, 4 ... 60 s backoff on its own
timer. A log with lines this build cannot decode starts a history epoch of its
own at start. far_e2e waits for the replacement owner to be listed before its
late push (the 403 it saw was the owner-machine check, not the epoch).
…pped and reconnected

Alive.due (core/desk.bend) decides on the web page, the native window's link
(bounded byte poll) and both phone apps (watchdog timers with the same
limit). The frame draws nothing and changes no state (laws hb_shows_nothing,
hubs_redraw_hb, hb_ignored_by_fold, alive_due). The kept-request helper for
the bot calls (Ui.rpc.keep.msg) rides along in client.bend.
…s msg key

The outbox replay carries the same msg, which is the entry's id; the hub
answers a repeat ok and stores, wakes and sends nothing (laws bots_say_once,
bots_post_once, bots_tell_once; test/bots_test.bend).
… own

The CLI writes one block per line, each numbered from 0, so text, thinking,
text collided on msg:0 and the id dedupe dropped the second. A line with a
uuid adds its first 8 characters to the id; stored ids are untouched (law
claude_text_ids_distinct, test/claude_cmds_test.bend).
…b; the project walk is cached 30 s

Hub.auto runs when the files arrive (AutoView). The Claude spawn stays on the
hub: a line for a thread with no process yet is dropped, so deferring the spawn
would lose messages; its walk (find through sh) is kept for 30 s per folder in
the user's cache instead.
…rite; a failed tail repair is reported, not dropped
…s end; retiring stops its watchers (laws watch_end_retired_silent, watch_retire_unwatches)
… again before the files are replaced and when the delayed restart runs
…s the command that was asked (laws watch_target_*)
…tion; a launch marker with no sign of a launch no longer suppresses it (test/tools/watch_scripts_test.ts)
i2cjak added 27 commits October 2, 2026 14:54
…ntinue is on; off clears it and starts nothing
…s a join frame as one frame of no bytes, as the hub charged it
…read there), a message naming $skills is prepared by the attach job, Codex's duplicate inline read is gone
An agent's unscoped child once grew to 47 GB inside the hub's own cgroup and
the kernel killed the whole service. The server now starts claude, codex and
grok through `systemd-run --user --scope` with MemoryMax (setting
agent.memory in GiB, clamped to 2 GiB..90% of RAM, default half; work_run
units get the same limit). A runaway is killed inside its scope; a signal exit
asks the scope for its Result and an oom-kill ends the turn with "agent ran
out of memory (limit N)". Without user systemd, or when the program is not
there, the agent starts as before. The prompt says jobs run under the limit.

Laws contain_*; test/contain_test.bend; e2e test/tools/contain_e2e.ts.
A streamed reply was laid out in full every frame. The window now splits it
at its last blank line outside a fenced block (client.bend's Live.at/head/
rest), lays the settled prefix out once in the thread's TLMemo (LiveMemo,
TL.live.m, kept by Lay.memo) and lays out only the tail each frame.

Laws live_split_*, live_prefix_exact (the prefix and the tail render to the
whole's nodes), tl_blocks_append and live_cat_exact (their layout is the
whole's); tests live_split_test, live_memo_test. The web and the phones
still lay the whole reply out.
@i2cjak
i2cjak merged commit 5dd1ded into main Oct 3, 2026
2 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant