Skip to content

fix(auth): explain expired environment credentials instead of offering a dead reconnect - #11462

Closed
gabrielsayers1 wants to merge 241 commits into
pingdotgg:mainfrom
goodbirdhq:phoenix/fix-agents-environment-not-connected
Closed

gabrielsayers1 wants to merge 241 commits into
pingdotgg:mainfrom
goodbirdhq:phoenix/fix-agents-environment-not-connected

Conversation

@gabrielsayers1

Copy link
Copy Markdown

Problem

A paired client's session lasts 30 days (DEFAULT_SESSION_TTL) and there is no renewal path, so every remote environment eventually hits expiry. When it does, the server knows exactly what happened — SessionStore.verify fails with SessionTokenExpiredError — but the wire error flattens that into invalid_credential. The client then shows a generic red "<label> is offline" banner whose Reconnect button re-sends the same dead credential forever.

The only visible difference from an ordinary network outage is the banner's colour, because ChatView renders the same title for supervisor phase offline and phase blocked. Diagnosing a real instance of this took reading server-side trace logs to discover the word "expired" had been thrown away.

Fix

Expiry now survives the trip to the client, and both clients act on it.

  • Contracts — EnvironmentAuthInvalidError carries an optional credentialExpired flag. It is an optional key sent only when true, following the existing dpopFailureReason precedent, so older clients keep decoding the payload and fall back to reason. Adding a new reason literal instead would have broken strict decoding on older clients.
  • Server — serverAuthCredentialExpired classifies the underlying cause, covering session and websocket expiry, and every failEnvironmentAuthInvalid call site passes it.
  • client-runtime — presentConnectionState exposes blockedReason, and a shared connectionNeedsPairing helper answers the question both surfaces were getting wrong: retrying cannot clear an authentication block. A DPoP proof hint is now suppressed for expiry, where it would misdirect.
  • Web — banner becomes "<label> needs pairing again" with a key icon, and the Reconnect button is gone. Connections remains as the way forward.
  • Mobile — same title change, and it stops telling the user it "will keep retrying automatically", which was actively false.

Desktop inherits the web change. The Connections panel improves for free via connectionStatusText.

Testing

vp test run across the touched scopes: 233 passing. Typecheck clean on contracts, client-runtime, server, web, and mobile; targeted lint and format clean.

Tests were written first and watched fail — server classification (expired vs. the revoked near-miss), the client message mapping, blockedReason exposure, and connectionNeedsPairing.

⚠️ Outstanding

Before/after screenshots are not attached, which this repo requires for UI changes. Capturing them means running the app in a browser, which needs maintainer sign-off — happy to produce them on request before merge.

Deliberately out of scope

  • No session renewal. There is no rotate/renew/refresh anywhere in server auth today; adding one is a new auth-protocol feature, not a bug fix. This PR makes expiry legible and recoverable, it does not stop the 30-day logout. Worth a follow-up, ideally with a warning before expiry while the client can still act.
  • retryNow does not re-sample connectivity (supervisor.ts), so Reconnect also silently fails while navigator.onLine is latched off after sleep. Real, separate concern, separate PR.

Model: Claude Opus 5 (1M context). Harness: Claude Code.

🤖 Generated with Claude Code

roughcoder and others added 30 commits August 11, 2026 18:39
* feat: agent sessions can spawn and orchestrate other sessions

A running agent session can now create sibling sessions on any configured
provider, exchange messages with them, and receive their completion reports
without polling. New sessions MCP toolkit (list_session_providers,
spawn_session, send_to_session, read_session, stop_session, post_report) on
the per-thread t3-code server; spawned threads get their own worktree by
default and carry a spawnedByThreadId link. A SessionSpawnReactor starts a
turn on the spawning thread when a child posts its report or hits a provider
error. Reports persist as a projection and render as cards; child threads
show a Spawned-by banner. Guardrails: direct-children-only scoping, 8-child /
3-deep caps, child permission mode capped at the parent's, and a Settings
toggle (enableSessionOrchestration) that applies to running sessions
immediately.

Also extracts the thread bootstrap macro from ws.ts into a shared
ThreadTurnBootstrap service, bumps @anthropic-ai/claude-agent-sdk to 0.3.226
(new SDK message types handled in ClaudeAdapter), and fixes an MCP gotcha
where an empty tool parameter struct made Claude Code drop the entire
server's toolset.

Built with Claude Fable 5 on Claude Code.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* feat: spawn_session accepts per-model provider options

list_session_providers now returns each model's option descriptors
(reasoning effort, context size, …) and spawn_session takes an optional
options array of {id, value} selections, validated against the chosen
model's descriptors before being carried into the thread's model selection.

Verified live: a Claude session spawned a Codex child with
reasoningEffort=low; the selection persisted on the spawned thread and the
completion report woke the parent.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
Renames the fork's user-visible identity from T3 Code to Phoenix and gives it a distinct OS-level
identity, so it installs and runs alongside upstream T3 Code rather than colliding with it.

Runtime identity diverges, source identity does not: state dir (~/.phoenix), env var (PHOENIX_HOME),
desktop userData, product name, URL scheme (phoenix://), bundle ID (com.goodbird.phoenix), Linux
desktop entry and WM class, systemd unit, ports (3873, 13873/5833), CLI binary and MCP server id all
change. The @t3tools package scope, t3-prefixed paths and internal symbols stay as upstream has them
so merges stay cheap. See docs/internals/branding.md.

Notably, the legacy-userData adoption step is removed entirely - keeping it would have pointed two
live applications at one directory - and PHOENIX_HOME deliberately has no T3CODE_HOME fallback,
unlike every other variable, because the base dir holds the SQLite database and auth state.

Also reworks CI for the fork. Upstream targets Blacksmith runners, which a fork without a Blacksmith
installation queues indefinitely rather than failing, so no check on this repository had ever
executed. Points ci.yml at GitHub-hosted runners, disables the nine workflows that need T3's
infrastructure (GitHub-side, so the files stay byte-identical to upstream), and adds phoenix-build.yml
to produce an unsigned macOS DMG per merge with a GitHub pre-release and Slack notification.

That first real CI run surfaced three defects that had been invisible on main, all from PR #1 landing
with zero executed checks: a missing `reports` field in a mobile test fixture, a missing
`spawnedByThreadId` in two ProjectionSnapshotQuery assertions, and an upstream image-compression test
that starves under parallel load. All three are fixed here.

Attribution to T3 Code is added to the README and Settings -> About; the original MIT copyright is
untouched.
The first Phoenix Build produced a DMG macOS refuses to open, reporting
"Phoenix (Alpha) is damaged and can't be opened".

Nothing was damaged. Passing CSC_IDENTITY_AUTO_DISCOVERY=false makes
electron-builder skip signing altogether, which leaves the bundle carrying
Electron's stock linker-signed signature - identifier "Electron", no sealed
resources - while the bundle around it has been rewritten. codesign --verify and
spctl both reject that state ("code has no resources but signature indicates
they must be present"), and on Apple silicon macOS refuses to launch it.

Supplies codesign's ad-hoc identity ("-") for unsigned macOS builds instead, and
stops disabling identity auto-discovery on macOS so electron-builder actually
runs the signing step. Ad-hoc signing needs no certificate and no Apple account,
and signs inside-out - helpers and frameworks before the outer app - which a
post-hoc `codesign --deep` does not do correctly.

Also corrects the release notes, which told users to clear the quarantine
attribute. That was the wrong remedy: the fault was an invalid signature, not
quarantine, so xattr would not have helped. The notes now describe the
notarisation prompt an ad-hoc signed build actually produces, and flag "is
damaged" as a bug to report rather than something to work around.
Swaps T3's wordmark for the feather-in-flame artwork across the production
channel: macOS/iOS/universal 1024 PNGs, the Windows ICO, and the web favicons.

Framing is cropped so the artwork fills 94% of the icon height rather than 81%.
The mark is tall and narrow (3080x4332 in a 5333 canvas), so the remaining
whitespace is at the sides and is a property of the shape, not the framing -
cropping further would leave no breathing room top and bottom.

The background stays white. A dark background also works, since the black quill
reads against the orange flame rather than against the backdrop, but white is
what was chosen for now and is a one-line change in icon.json to revisit.

The assets were generated directly rather than through `pnpm icons:export`:
that pipeline needs Icon Composer 2.x for design generation 26, and Xcode 26.4.1
ships version 1.4. The .icon project is updated to reference the new artwork so
the pipeline stays correct for whoever has a compatible Icon Composer, but the
committed PNGs are what the desktop build actually consumes - stageMacIcons
feeds the 1024 PNG through sips and iconutil at package time. `icons:check` is
not run in CI, so the two will not conflict.

Development and nightly channels still carry the T3 artwork, which incidentally
makes local dev builds easy to tell apart from a release build in the Dock.
First launch of the packaged app asked to unlock "t3code Safe Storage", not a
Phoenix-owned item. Electron derives app.getName() from the staged app's
package.json name, and macOS safeStorage keys its Keychain entry off that
("<name> Safe Storage"). The staged name was still "t3code", so Phoenix and
upstream T3 Code encrypted their secrets against a single shared keychain item -
the exact class of collision the separate runtime identity exists to prevent.

Sets the staged name to "phoenix". This is distinct from the workspace package
name in apps/server/package.json, which stays "t3" because it never reaches the
OS.

docs/internals/branding.md claimed the Keychain item followed productName and
would therefore be "Phoenix (Alpha) Safe Storage". That was wrong on both counts
and had this identifier filed under source identity, which is why it was missed.
Corrected.

Also fixes the release notes, which told users to Control-click and choose Open.
macOS 15 removed that bypass for un-notarised apps, so System Settings is now
the only route; the notes describe that flow and the actual dialog text.
The desktop sidebar still read "T3 Code" after the rebrand. The name was never
a string: SidebarBrand rendered an inline SVG of the "T3" glyph followed by a
separate hardcoded "Code" literal, so no search for "T3 Code" could ever have
matched it. The mobile home header and compact title did the same.

Drives the desktop sidebar from APP_BASE_NAME so it follows the branding
constant rather than hardcoded artwork, and uses the name directly on mobile.
Deletes T3Wordmark, which has no remaining callers.

This is the third rebrand miss found by running the app rather than by reading
it, after the broken signature and the shared keychain item. All three were
invisible to grep: artwork, an OS-derived identifier, and a split literal.
A running Phoenix install bound port 3774 while 3873 sat free: it had tried
3773, found upstream T3 Code already there, and scanned forward. Adjacent ports
look harmless, but the order decides the outcome - start Phoenix first and it
takes 3773, pushing T3 Code off its own port.

ServerConfig.DEFAULT_PORT was moved to 3873 in the rebrand, but the desktop
keeps its own constant and never consulted it, so the change had no effect on
the packaged app. packages/ssh had a third copy for remote tunnels.

Points DEFAULT_DESKTOP_BACKEND_PORT and DEFAULT_REMOTE_PORT at 3873, and
corrects the scan-range comment in the server's auth notes.

Found by inspecting a running install rather than by reading the code - the same
way the broken signature, the shared keychain item and the T3 wordmark surfaced.
Normalize the feather source for Icon Composer and regenerate the production macOS artwork with a rounded, inset body so Dock rendering no longer exposes an opaque square.

Constraint: classic macOS icons require an inset body with transparent canvas around it.

Rejected: patch the installed signed app bundle | release assets must come from tracked source.

Confidence: high

Scope-risk: narrow

Directive: preserve the transparent macOS safe area in future production icon exports.

Tested: native Icon Composer iOS/macOS renders; iconutil ICNS packaging; extracted 1024px rendition; git diff --check.

Not-tested: installed release bundle; no release was requested.
Sync the fork with pingdotgg/t3code through 63e6fae.

Notable upstream additions: Open VSX theme search (pingdotgg#5654), OKLCH theme
palettes (pingdotgg#6036), hourly past-24-hour usage view (pingdotgg#6170), mobile thread
title regeneration (pingdotgg#6253), and Windows/Azure DevOps/GitLab fixes.

One conflict, in apps/web/src/components/settings/ThemeEditorPanel.tsx.
Upstream pingdotgg#6183 makes environment artwork theme aware and removes the
per-theme opt-in toggle; our rebrand had only relabelled that toggle's
copy to "Show Phoenix environment artwork". Resolved in favour of
upstream, since the surrounding sidebarArtwork state it referenced is
gone. The themePalette data model upstream retains is untouched.

Verified: pnpm typecheck (no errors), pnpm test (2381 passed, 7 skipped).

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
… documented limits (#9)

* feat(sessions): document messageLimit bounds, add structured report fields, fix stop_session null result

- read_session/post_report tool descriptions now state messageLimit's
  0-20 range, 5 default, and 16,384-char per-message truncation
  (the schema already enforced min/max; only the description was missing it).
- post_report accepts optional findings/validation/recommendation/
  completionPercent fields alongside the markdown summary, threaded through
  the command, event, and a new JSON column (migration 043) so read_session
  can surface them. All additive, no breaking changes.
- stop_session now returns { threadId, status } instead of null, which was
  failing MCP client schema validation ("expected record, received null")
  even though the stop succeeded server-side.
- spawn_session's description documents the SESSION_SPAWN_MAX_CHILDREN cap,
  and the spawn-limit error message now correctly says archiving (not
  stopping) is what frees a slot.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* fix(sessions): bound structured report payload size, degrade gracefully on bad JSON

Addresses PR #9 review feedback:

- findings/validation.performed/validation.gaps arrays were unbounded while
  scalar fields were capped, letting a huge findings array store and return
  wholesale into a parent's context via read_session. Cap findings to 100
  entries, performed/gaps to 50 each, and add a 32KB combined encoded-size
  cap across findings/validation/recommendation/completionPercent as
  defense in depth.
- A malformed structured_json blob previously failed the whole DB row via
  strict schema decoding, which would have blocked read_session and thread
  detail hydration. Decode it leniently instead (new
  decodeStructuredReportFields helper): parse/schema failures log a warning
  and are treated as absent, never blocking report access.
- Services/ProjectionThreadReports.ts's completionPercent was a plain
  Schema.Int with no 0-100 bound; it now reuses SessionReportStructured's
  fields directly so persistence-layer bounds can't drift from the
  contract.
- encodeStructured now writes SQL NULL instead of "{}" when there's nothing
  to store, keeping "no structured data" and "empty structured data"
  distinguishable in the column.
- Renamed handlers.test.ts to tools.test.ts (it only exercised
  ReadSessionInput/PostReportInput schema decoding, not any handler), and
  added a focused decodeStructuredReportFields.test.ts plus size/array-cap
  cases to tools.test.ts.

Migration numbering and #8 untouched, per review instructions.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
* feat(sessions): let spawned agents inspect a specified revision

Constraint: Keep checkout creation inside the existing bootstrap rollback boundary.\nRejected: Ad-hoc shell commands in the MCP handler | GitWorkflowService already exposes the Git driver seam.\nConfidence: high\nScope-risk: narrow\nDirective: Worktree cleanup and reclamation remain a separate lifecycle concern.\nTested: vp test run packages/contracts/src/sessionOrchestration.test.ts apps/server/src/mcp/toolkits/sessions/handlers.test.ts; npm run typecheck.\nNot-tested: Live provider/session spawn against a remote pull request.

* fix(sessions): keep checkout metadata best effort

Constraint: Session reports and controls must survive missing or unreadable worktrees.\nRejected: Making checkout metadata a required read/spawn prerequisite | optional fields must not orphan live sessions.\nConfidence: high\nScope-risk: narrow\nDirective: PR-ref portability beyond GitHub remains follow-up work.\nTested: vp test run apps/server/src/mcp/toolkits/sessions/handlers.test.ts apps/server/src/orchestration/ThreadTurnBootstrap.test.ts packages/contracts/src/sessionOrchestration.test.ts; npm run typecheck.\nNot-tested: Live remote fetch and provider-session spawn.

* test(sessions): type checkout failure doubles accurately

Constraint: Tests must match nominal Git workflow error and status contracts.\nRejected: Untyped Error and partial result stubs | root tsgo rejects the incompatible channels.\nConfidence: high\nScope-risk: narrow\nDirective: Keep optional checkout enrichment non-blocking.\nTested: npm run typecheck.\nNot-tested: Focused Vitest suite not rerun; this change is test typing and assertions only.

* fix(sessions): use typed checkout failure channels

Constraint: Effect code must use platform services and tagged failure values.\nRejected: Node path import and generic Error failures | root tsgo treats them as errors.\nConfidence: high\nScope-risk: narrow\nDirective: Preserve the fetch-versus-missing-ref distinction through typed errors.\nTested: npm run typecheck; vp run --last-details (all 15 tasks succeeded, exit 0).\nNot-tested: Focused Vitest suite not rerun.

* test(server): resolve bootstrap revisions in seam harness

Constraint: Bootstrap now resolves a commit before creating a worktree.\nRejected: Leaving the driver mock implicit | seam tests must exercise the production dependency path.\nConfidence: high\nScope-risk: narrow\nDirective: Keep revision stubs transparent so base-ref assertions remain meaningful.\nTested: vp test run apps/server/src/server.test.ts -t five bootstrap seam cases; vp test run apps/server/src/mcp/toolkits/sessions/handlers.test.ts apps/server/src/orchestration/ThreadTurnBootstrap.test.ts packages/contracts/src/sessionOrchestration.test.ts; npm run typecheck (exit 0).\nNot-tested: Full t3 server suite.
…#6)

* feat: let spawned sessions stop gracefully

Constraint: Stop requests must remain command-to-event-to-projection flows.\nRejected: A parallel stop-state store | it would bypass durable session projection.\nConfidence: medium\nScope-risk: moderate\nDirective: Keep MCP tool input schemas explicitly typed; no empty structs.\nTested: targeted immediate-stop reactor test; focused migration, contract, and decider tests.\nNot-tested: full typecheck remains blocked by the machine's Node/dependency setup and the resource cap.

* fix: preserve graceful stop invariants

Persist grace deadline episodes so restart recovery and a resumed session cannot be stopped by stale work.

Constraint: graceful stops must survive reactor restarts without blocking ordinary session resume.
Rejected: in-memory daemon-only deadline | lost on restart and could stop a new session.
Confidence: high
Scope-risk: moderate
Directive: keep grace notices explicitly tagged; do not gate normal turn starts on stopped status.
Tested: focused graceful deadline and stopped-session resume tests; decider, sessions handler, and migration tests; npm run typecheck (exit 0).
Not-tested: provider-specific steer behavior.

* fix: keep rebased graceful stop typed

Constraint: session stop must coexist with the merged report and checkout contracts.
Rejected: retaining pre-rebase type assumptions | stale contract and Effect diagnostics failed typecheck.
Confidence: high
Scope-risk: narrow
Directive: keep stop-session schema fixtures branded and time through Effect DateTime.
Tested: npm run typecheck -- --pretty false; targeted session, decider, reactor, and migration tests.
Not-tested: provider-specific steer behavior.

* test: cover graceful stop session defaults

Constraint: projected session snapshots include stop-audit and graceful-deadline defaults.
Rejected: partial fixture update | snapshot hydration compares the complete session shape.
Confidence: high
Scope-risk: narrow
Directive: update complete session fixtures when projection fields are added.
Tested: targeted ProjectionSnapshotQuery, sessions, decider, reactor, and migration tests (95 passed).
Not-tested: full typecheck; prior gate run exited 0.
#5)

* feat(server): make session messages queue safely

Persist busy-child deliveries as FIFO queued turn starts and release them at provider session boundaries. Interrupt mode requests the existing provider interrupt before the queued replacement is released.

Constraint: send_to_session must remain event-sourced and provider-neutral.

Rejected: provider-level notify steering | adapter semantics are inconsistent and need an explicit capability contract.

Confidence: high

Scope-risk: moderate

Directive: add notify only after every provider has an explicit steering or fallback contract.

Tested: pnpm exec vp test run apps/server/src/mcp/toolkits/sessions/handlers.test.ts apps/server/src/orchestration/decider.settled.test.ts (15 passed); npm run typecheck under Node 24 found one startup error-channel leak, fixed afterward.

Not-tested: full typecheck was not rerun due the session's one-run resource constraint; no browser or full-suite validation.

* fix(server): keep queued session delivery live

Make queued delivery release idempotent, recover stale persisted sessions only after checking live provider bindings, cancel terminal queues, and bound interrupt fallback with parent-visible cancellation.

Constraint: queued delivery must remain event-sourced and safe across restart, provider failure, and concurrent release signals.

Rejected: unconditional stop-then-send fallback | provider exit can race terminal queue cancellation and resurrect stopped work.

Confidence: high

Scope-risk: moderate

Directive: preserve queued IDs in the command read model when changing turn projections.

Tested: pnpm exec vp test run apps/server/src/orchestration/Layers/SessionSpawnReactor.test.ts apps/server/src/orchestration/decider.settled.test.ts apps/server/src/mcp/toolkits/sessions/handlers.test.ts apps/server/src/orchestration/projector.test.ts apps/server/src/orchestration/Layers/ProjectionSnapshotQuery.test.ts (50 passed); npm run typecheck under Node 24 found four localized typing issues, fixed afterward.

Not-tested: typecheck was not rerun due the one-run resource constraint; no full suite or browser validation.

* fix(server): recheck queued session liveness

Move provider liveness observation into the serialized recovery decision and make post-dispatch acknowledgement uncertainty explicit.

Constraint: Shared-machine validation permits one typecheck run and targeted tests only.
Rejected: Two-observation failure detection | a decision-time in-memory provider check closes the stale snapshot window directly.
Confidence: high
Scope-risk: narrow
Directive: Keep provider liveness checks inside serialized recovery decisions.
Tested: pnpm exec vp test run apps/server/src/orchestration/Layers/SessionSpawnReactor.test.ts apps/server/src/mcp/toolkits/sessions/handlers.test.ts apps/server/src/orchestration/decider.settled.test.ts (21 passed, exit 0).
Not-tested: npm run typecheck was not rerun after fixing its two exact diagnostics; its single permitted run exited 1.

* fix(server): document recovery binding invariant

Record the provider-directory assumption at the recovery decision and keep recovery sweep interruption out of the error channel.

Constraint: Recovery may force-release only when directory absence is authoritative.
Rejected: Optional cleanup refactors | closing the reviewed invariant and compiler issue keeps this commit narrow.
Confidence: high
Scope-risk: narrow
Directive: Revisit the recovery failure detector if provider directory absence can become transient during rebind.
Tested: Node 24 npm run typecheck (exit 0); targeted SessionSpawnReactor test (6 passed, exit 0); git diff --check.
Not-tested: Full test suite and browser validation were not run.

* fix(server): preserve graceful stop steering

Keep grace-stop notices immediate while ordinary busy-session messages queue, and make provider reactor fixtures model a completed turn before follow-up commands.

Constraint: Rebase must preserve graceful-stop deadlines and queued session delivery semantics.
Rejected: Queueing grace notices | it can hide the stop deadline until the turn it must stop has already ended.
Confidence: high
Scope-risk: narrow
Directive: Internal in-turn steering commands must opt out explicitly if future delivery defaults queue busy sessions.
Tested: Node 24 npm run typecheck (exit 0); targeted provider/handlers tests (53 passed, exit 0); targeted decider/session reactor/snapshot tests (39 passed, exit 0); git diff --check.
Not-tested: Full repository test suite and browser validation were not run.

* test(server): preserve pending steer coverage

Exercise conflicting provider turn acceptance through the intentional immediate grace-notice steer now that ordinary busy-session messages queue by default.

Constraint: The ingestion test must create a pending turn start before asserting conflicting turn acceptance.
Rejected: Restoring implicit busy-session steering | it would violate send_to_session's queue-by-default contract.
Confidence: high
Scope-risk: narrow
Directive: Pending-turn ingestion tests must use an explicit immediate-start path when the thread is already running.
Tested: Node 24 npm run typecheck (exit 0); targeted ProviderRuntimeIngestion, ProviderCommandReactor, SessionSpawnReactor, and decider suites (112 passed, exit 0); single regression test (exit 0); git diff --check.
Not-tested: Full repository suite and browser validation were not rerun locally.
…e_session with worktree cleanup (#10)

* feat(orchestration): synthesize terminal reports and add settle_session

A spawned session that died without calling post_report left its parent
with nothing to act on: a `stopped` child produced no notification at all,
and an errored one produced a bare error message with no record of what it
had done. Separately, nothing in the server ever reclaimed a spawned
worktree, so a long orchestration run leaked a directory per child.

Synthetic terminal reports: when a spawned child's session reaches
`stopped` or `error` with no report posted, SessionSpawnReactor dispatches
an ordinary `thread.report.post` carrying the termination reason, the last
tool activity and assistant message, and an explicit "work is likely
unfinished" warning. Everything downstream — projection, the report card,
the parent wake-up — behaves as it does for an agent-posted report, which
is how `stopped` now notifies the parent at all. `SessionReport.origin`
("agent" | "system", migration 043) keeps a synthesized report from ever
reading as the child's own claim, in the parent's notification and in the
UI badge.

settle_session: the parent explicitly settles a child rather than sessions
settling themselves. It refuses a starting/running child with an
actionable message instead of stopping it, and keeps the decider's guards
for open approvals and queued turns. `cleanupWorktree: true` permanently
deletes the child's worktree and its temporary `t3code/…` branch — refused,
with the specific dirty files and unpushed commit count, unless the work is
committed and pushed or `force: true` is passed. Branches Phoenix did not
create are kept and reported. The result always names what was removed and
what was kept, and read_session now exposes the child's worktree path so a
deleted worktree stops being advertised.

Validation: `npm run typecheck` clean; new handler and reactor tests plus
the decider report tests pass.

* fix(orchestration): stop a live session on settle, mark reports only once posted

Addresses code review on PR #10.

[major] The terminal-report dedup set was marked before the dispatch that
posts the report. processEventSafely swallows dispatch failures, and a
terminated session emits no further status transition, so a transient
failure branded the episode "handled" with nothing persisted and the parent
was never told — worse than the old behaviour, which at least never
pretended. The episode is now marked only once the report is actually
posted, the persisted `reports` check remains the real duplicate guard, and
the dispatch retries twice before giving up, since this is the last chance
to reach the parent.

[major] settle_session now stops a still-alive session instead of leaving a
`ready` provider process running behind a settled thread. A turn in flight
(starting/running) is still refused — interrupting live work stays a
deliberate stop_session — but an idle-but-alive child is stopped as part of
settling, because settling is the parent declaring it finished. Order is
stop -> settle -> assess -> remove, which also closes the window the
reviewer flagged: the dirty/unpushed check now runs after the process is
gone, so nothing can write to the worktree between the check and the
removal. If the session does not reach "stopped" within the timeout the
thread is still settled but cleanup is withheld and the caller told.

[minor] Documented that deleteRef is a trusted-caller primitive: the
branch-safety decision lives in decideBranchCleanup, not in the driver.

Tests: both majors were in wiring that only pure-helper tests covered.
Added SessionSpawnReactor.wiring.test.ts driving the real reactor against a
stub engine (verified failing against the old ordering: 3 dispatch attempts
instead of 6), and settleSession.test.ts driving the real handler against
stub services, asserting the stop/settle/inspect/delete call order.

Validation: `npm run typecheck` exit code 0; 34 tests across the five
touched files pass; `vp lint` clean.

* fix(contracts): supply required stop-audit fields in the checkout result test

Pre-existing on main, surfaced by rebasing onto it: #4's ReadSessionResult
decode fixture does not supply the stoppedBy / stopRequestedAt / stopReason /
interruptedToolCall / lastCompletedOperation fields that #6 made required, so
`packages/contracts/src/sessionOrchestration.test.ts` fails at origin/main
with "SchemaError: Missing key at [stoppedBy]". Verified by running the test
against main's own unmodified sessionOrchestration.ts.

Not caused by this branch, but it would land on top of it, so it is fixed
here rather than left for the merge.
… sibling access (#8)

Rebased onto main after #4/#5/#6/#9/#10 merged; carries the reconciliation:

- Envelope delivery applies to agent AND system-origin synthetic reports
  (same report-posted path); formatReportMessage keeps #10's origin-aware
  lead and delivers >1KB summaries as an envelope with a read_report hint.
- SessionReportEnvelope carries #9's structured data compactly:
  recommendation and completionPercent whole, findings/validation as
  counts; and #10's origin so a synthesized epitaph is visible at a
  glance. read_report returns the full findings/validation arrays with
  every page (bounded by the 32KB structured cap) plus origin.
- findByReportId now selects abstract, structured_json, and origin and
  maps through the shared mapReportRow, so no column can be silently
  dropped by one read path; the ProjectionSnapshotQuery Struct.pick
  allowlist gained abstract.
- Migration renumbered 043 -> 046 (043 structured, 044 stop audit, 045
  origin landed first).
- read_report docs record UTF-16 code-unit paging (pages can run one unit
  short or long at surrogate boundaries) and the unguessable-UUID timing
  assumption behind the reportId-only lookup path.

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
Expose the in-memory background-liveness and plan-progress state (already
used by the sidebar) over MCP without starting a turn or disturbing the
child: a new ping_session tool returns session status, settlement, last
provider activity timestamp, current background activity, plan progress,
whether a report has landed, and a snippet of the last assistant message.
read_session gains the same lastActivityAt/currentActivity fields
(additive, optional) for parity.

hasReport and lastAssistantMessage are served by two purpose-built,
bounded ProjectionSnapshotQuery reads (getThreadHasReport,
getLastAssistantMessage) rather than a full/windowed thread detail
snapshot, so ping_session stays cheap regardless of a child's history,
and the last-assistant-message lookup is correct even when the child's
newest turn happens to be user-only. All optional enrichment (provider
session directory lookup, report existence, last message) is
best-effort: a lookup failure degrades to null/false rather than
failing the call.

Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
…h merged session features (#11)

Assert graceful-stop audit defaults explicitly and run Effect-based tests through @effect/vitest so main's release gate matches the merged runtime behavior.

Constraint: Restore main CI with one narrow test-only PR.
Rejected: Removing stop audit metadata | the client command intentionally propagates user stop attribution.
Confidence: high
Scope-risk: narrow
Directive: Use @effect/vitest for tests that execute Effect values; do not call manual runtimes.
Tested: vp lint (exit 0); Node 24 npm run typecheck (exit 0); focused four-file tests (102 passed, exit 0); root vp run test after ensuring Electron runtime (2485 passed, 7 skipped, exit 0); git diff --check.
Not-tested: Browser and mobile device integration were not run.
* feat(contracts): add sidebarSessionHierarchyEnabled client setting

A rendering mode over the spawnedByThreadId links the server already
stores, so it lives with the other sidebar client settings rather than
in ServerSettings: nothing about it is server-authoritative.

Defaults off. Turning it on changes how every existing thread list is
laid out, so it opts in rather than surprising users on upgrade.

The desktop client-settings fixture builds a complete ClientSettings
literal, so every new key has to be added there too.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* feat(web): build the sidebar session tree from spawnedByThreadId

buildSidebarThreadHierarchy nests threads under the session that spawned
them, recursively and to any depth. Sibling order is the incoming order:
the caller has already applied the section's sort, and hierarchy mode
re-parents rows without reordering them.

Two invariants matter more than the nesting itself, because both failure
modes hide work from the user:

- Every input thread is emitted exactly once. A parent that isn't in the
  list (archived, snoozed, settled into another section, filtered out by
  project scope) leaves its child at top level instead of dropping it,
  and threads caught in a parent cycle are emitted flat for the same
  reason.
- Parent links are matched per environment. The link is a bare thread id,
  so two environments can hold the same id without being related.

Indentation caps at four levels while depth keeps counting, so a deep
tree stops stealing horizontal room from titles but still reads as
nested.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* feat(web): render nested sessions as indented, compact rows

A spawned thread renders under its parent, indented one step per level
with a hairline rail down the gutter so the nesting is followable
without counting pixels.

Nested rows drop the branch line and shrink from 4.875rem to 3.375rem.
Children still have their own worktrees — the branch goes for density,
because a tree of full-height cards buries the parent it belongs to. The
signal that line carried moves rather than disappearing: the PR number
and terminal indicator ride up beside the status, since the PR is
usually the whole point of a spawned session.

Hierarchy applies to the inbox only. The other blocks each carry an
ordering that nesting would fight: pinned rows are a hand-dragged
arrangement, the snoozed shelf sorts by what wakes next, and the settled
tail is paginated history where a parent and its children could straddle
a page boundary.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* feat(web): add the sidebar hierarchy toggle under session orchestration

The switch renders only while session orchestration is on: with it off
nothing spawns children, so the mode would have nothing to nest and the
row would be a dead end.

The gate is on the settings row, not on rendering. An existing tree
keeps nesting if orchestration is later turned off, because the nesting
is history rather than a live permission.

Registered with the settings search index and the dirty/reset-all lists
alongside the other client settings.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* docs: document sidebar session hierarchy

Covers what nests, the four-level indent cap, why child rows are shorter
(and that children keep their own worktrees), which sections hierarchy
applies to, and that a child outlives a parent that leaves the inbox.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
… 047) (#18)

* feat(persistence): reserve queued delivery receipts

Constraint: Migration 047 must land before migration 048 in the merge train.
Rejected: Bundling schema reservation with delivery behavior | migration ordering blocks independent PRs.
Confidence: high
Scope-risk: narrow
Directive: Keep migration 047 limited to durable receipt columns and its query index.
Tested: migration test pending.
Not-tested: full suite.

* fix(persistence): index receipt ordering

Constraint: Migration 047's receipt index must serve the bounded query order.
Rejected: State-first ordering | it cannot support the receipt timestamp expression.
Confidence: high
Scope-risk: narrow
Tested: migration test pending.
…15)

* feat(contracts): add SessionUsageSnapshot for session budgeting

Best-effort token/turn accounting so an orchestrating parent can budget
instead of flying blind: threaded through PingSessionResult, SessionReport,
SessionReportEnvelope, and ReadReportResult. Individual fields are optional
since a provider may not report tokens; deliberately no cost estimate, as
price tables go stale and tokens are the stable currency.

* feat(server): capture usage snapshots in ping_session and posted reports

Reuses the existing context-window usage projection (the same
projection_thread_activities rows the web client's context-window meter
reads) plus two bounded, single-purpose queries (getLatestUsageActivity,
getThreadTurnCount) rather than adding new provider plumbing. Usage is
computed server-side at ping/report-post time and stored alongside the
existing structured report fields, so no migration is needed. Both bounded
reads degrade to "field omitted" on failure, matching ping_session's
existing enrichment contract.

* docs(session-orchestration): document the usage snapshot

Explains where the usage numbers come from (the existing context-window
projection, not new provider plumbing), why no cost estimate is included,
and the report-post-time capture semantics.

* fix(sessions): make usage snapshot fields honest, fix turnCount, harden clocks

Address REQUEST-CHANGES review feedback on PR #15:

- Rename inputTokens/outputTokens to lastTurnInputTokens/lastTurnOutputTokens
  everywhere: both provider adapters report only the most recent turn's
  counts at this level, never a session accumulation, and the old names
  read as spend across the whole session. totalTokens is now sourced only
  from a provider's own cumulative counter (totalProcessedTokens) and
  omitted rather than backfilled from context-window occupancy
  (usedTokens), which is bounded by the window and drops after compaction
  — the opposite of a monotonic spend number.
- getThreadTurnCount now filters turn_id IS NOT NULL, excluding pending
  turn-start placeholders and queued/interrupting rows that share the
  table but are not turns.
- Guard Date.parse failures in elapsedMsSince/lastTurnDurationMsFromTurn
  with Number.isFinite so a bad timestamp degrades to 0/null instead of
  NaN, which would otherwise void the entire structured_json blob
  (findings/validation/recommendation/completionPercent alongside usage)
  through the lenient decoder, or fail ping_session's schema encode.
- Document cross-provider incomparability of the last-turn token fields
  (Claude folds cache tokens into input; Codex reports them separately).
…hains (#14)

* feat(contracts): optional report amendment links on reports and envelopes

post_report gains supersedesReportId so a session can amend an account it
already gave — the case that motivated this is a queued instruction landing
after the child reported, leaving a stale report as the record.

Both fields are optional on SessionReport and on the thread.report.post
command / thread.report-posted payload, so every already-persisted report
event replays unchanged. supersededByReportId is deliberately absent from the
event: at post time no superseding report exists yet, so it is derived on read
paths instead.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* feat(server): persist report amendments and derive the reverse link

Migration 048 adds projection_thread_reports.supersedes_report_id plus the
index every read path uses to resolve the other direction. Only the forward
link is stored: an amendment never rewrites the row it supersedes, so the
projection stays append-only and "who superseded me" is a correlated subquery
rather than a flag that could drift out of sync.

Migration id 48 leaves a gap at 47, which a sibling wave-2 slice reserves;
ids only have to be ordered, not contiguous.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* feat(sessions): amend reports via post_report supersedesReportId

post_report validates the reference before dispatching: it must name a report
on the calling thread. A report is a session's account of its own work, so
amending another session's report is refused — with the same message as an
unknown id, so the denial cannot double as a probe for report ids elsewhere.
A dangling link would be worse than a rejection: no reader could follow it.

Both ends travel outward. An amending report's parent notification leads with
"AMENDED report (supersedes ...)" before the summary, because a parent that
already acted on the superseded report has to see that first. read_report on a
superseded report still serves its body, plus supersededByReportId and a
supersededNotice sentence — a caller paging an old report cannot be relied on
to notice a field it was not looking for.

The spawned-session instructions now state the rule directly: an instruction
arriving after you reported means post an amending report, never claim
retroactive compliance.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* docs(internals): record how report amendment works

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(sessions): keep amendment chains linear, enforced in the decider

Review found A→B and A→C was possible, leaving reverse navigation and
latest-report selection free to disagree about which report is current.
Superseding an already-superseded report is now refused and the caller is
handed the head of the chain, which also settles the concurrent-fork race:
the second writer loses with an error naming where to re-attach.

The authoritative check lives in the decider, not the toolkit. Handler-level
validation cannot prevent a fork — two amendments can both pass their
pre-checks before either dispatches — whereas the decider runs against the
folded read model serialized with command processing, and also covers
internally dispatched report posts that never reach the toolkit. The toolkit
keeps a pre-check purely for error quality, and re-reads the chain when a
dispatch is rejected so a race loser gets the same structured error rather
than a generic dispatch failure. Both sides share one implementation and one
wording so they cannot disagree.

Tests: fork rejection including the race shape, chain-head reporting through
a longer chain, a one-turn reactor run asserting AMENDED delivery reaches the
parent (not just the formatter), and synthesized-report resurrection — a
Phoenix terminal report superseded by the agent's own account after resume.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
…-proven branch deletion (#17)

* feat(git): add a per-repository lock for git worktree mutations

Git allows one writer per repository: concurrent `git worktree remove` runs on
one repo contend for .git/index.lock, and the losers block until the caller's
own timeout kills them rather than queueing. This is one Effect Semaphore per
repository root so callers can serialize the mutations that take that lock.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* feat(contracts): branch cleanup, settle warnings, and git-hygiene errors

settle_session gains `cleanupBranch` (delete a branch Phoenix did not create,
against a merge proof), a `branchProof` on the worktree outcome, and a
`warning` on the result for a settle whose child process outlived its stop.

Two structured errors carry what prose cannot: which leg of the merge proof
failed with the SHAs that disagree, and which git lock file is in the way plus
the remedy. `ChangeRequest.headRefOid` is the merged commit the proof compares
against; optional, since not every host reports one.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* feat(sessions): serialize worktree cleanup and prove branch merges before deleting

Settling eight children at once with cleanupWorktree made every cleanup fail:
concurrent `git worktree remove` calls fought over .git/index.lock until each
timed out, and the one that ran alone was the only one that worked. A timed-out
removal then left a zero-byte index.lock that blocked every later git command
on the repository.

- Worktree removal and branch deletion now run inside GitRepositoryLock, keyed
  by the project root, so parallel settles queue instead of racing. Sessions
  still settle independently of cleanup.
- A git failure naming a lock file answers with SessionOrchestrationGitLockError:
  the path, its age, whether it matches the conservative stale heuristic (empty
  and older than 60s), and the remedy. Phoenix never deletes the lock itself —
  nothing in this process can prove no live git owns it.
- cleanupBranch deletes a user's branch only when local head == remote head ==
  the head commit of a merged PR. `git branch --merged` is useless here: this
  repo squash-merges, so a merged branch is never an ancestor of main. The
  proof runs before anything is destroyed, so a refusal costs nothing.
- A settle whose stop wait times out no longer succeeds silently; it names the
  provider and the status the session was last seen in.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* docs(sessions): worktree cleanup git hygiene

What changed and why: one cleanup per repository at a time, a leftover lock
reported rather than removed (with the heuristic and why Phoenix will not
delete it), the squash-merge trap that makes `git branch --merged` the wrong
proof, and the settle warning for a process that outlived its stop.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(sessions): make the lock error reachable, the lock key canonical, and the merge proof race-free

Review findings on #17, all four of which were the difference between code
that looks right and code that works.

- The structured lock error could never fire in production: the driver threw
  away git's stderr and substituted a fixed detail string, so only a test's
  injected message ever matched. GitCommandError now carries a bounded
  stderrExcerpt (tail-first — the fatal line is last), and the driver attaches
  it to every non-zero exit. Proven end to end against a real held ref lock:
  real repo → real git failure → structured error naming the real lock file.
- The semaphore was keyed on a trimmed path, so a symlinked alias or a linked
  worktree of the same repository got a *different* permit and serialized
  nothing. The key is now realpath + the repository's common git directory,
  with tests for both aliases (they fail against the old keying).
- The merge proof was computed before the semaphore and acted on inside it, so
  a branch updated while a cleanup queued would be deleted on a stale proof.
  The two heads are re-read inside the critical section immediately before the
  delete; a branch that moved is kept and reported instead.
- Lock matching keys on the .lock path rather than on git's English, and a
  branch delete that fails on a lock now raises the same structured error
  instead of swallowing it into a detail string — after clearing the thread's
  worktreePath, since the directory really is gone by then.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(vcs): redact credentials from git error excerpts, and delete proven branches by compare-and-swap

Round-3 review on #17.

- The stderr excerpt I added last round crossed RPC unredacted, and git echoes
  remote URLs that can carry their own credentials
  (https://x-access-token:ghs_...@host/o/r). Userinfo, query strings,
  fragments, and known token shapes are stripped before the excerpt is
  attached — and before it is truncated, so nothing survives on a boundary.
  Scheme, host, and path stay: which remote failed is the diagnostic value.
  This restores an invariant an existing driver test already asserted for
  args and raw stderr, and which the excerpt had quietly reversed.
- The in-lock re-proof is only atomic against writers in this process; a
  terminal or a second Phoenix can still move the ref between the check and
  the delete. deleteRef now takes an expectedSha and routes through
  `git update-ref -d <ref> <old>`, so git itself refuses a moved ref under its
  own lock. Temporary t3code branches keep the plain force delete: no proof,
  nothing to hold git to.
- A cleanup that removes the worktree but keeps the branch now reports a
  structured worktree.branchRefusal (branch, reason, SHAs, the expected
  commit) sharing its reason vocabulary with the BranchNotMerged error,
  instead of folding a partial success into a prose detail.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(sessions): refuse to delete a branch another worktree still has checked out

Round-4 review on #17. The compare-and-swap closed the ref race but opened a
hole: `git update-ref -d` is plumbing and, unlike `git branch -d`, deletes a
branch another linked worktree is sitting on — leaving that worktree's HEAD
pointing at nothing. Removing our own worktree first is not enough, because
the other one can be a worktree a user made by hand that this process has
never heard of and the repository lock does not cover.

So inside the critical section, before the CAS, the worktree list is consulted
and a branch still held anywhere is kept with a structured refusal naming the
conflicting path. An unreadable worktree list fails closed for the same
reason. The CAS stays as the final atomicity layer.

The driver test proves the hazard rather than assuming it: `git branch -d`
refuses a checked-out branch, plumbing deletes it. The settle test is the
real regression the reviewer asked for — a real repository with a real second
worktree, refused; the worktree removed; deleted.

Also corrects a comment that claimed argv never carries credentials: clone,
fetch, and push take the remote URL as an argument, which is exactly why
redaction is written against URL shapes rather than a command allowlist.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(vcs): delete branches with porcelain, not compare-and-swap

Design reversal, and the reasoning I gave last round was wrong. I argued the
compare-and-swap bounded the worktree TOCTOU. It does not bound it at all: a
checkout does not move the branch's OID, so `update-ref -d <ref> <sha>`
succeeds while another worktree holds the branch and leaves that worktree's
HEAD dangling.

The two remaining races are not equally survivable, and that decides the
mechanism. A checkout we get wrong corrupts someone else's directory and no
reflog repairs it. A ref move we miss costs a ref pointer the reflog still
holds. `git branch -D` refuses a checked-out branch atomically at delete time
and cannot see a ref move; the compare-and-swap is exactly the reverse. So:

- deleteRef is porcelain again, and the expectedSha parameter is removed
  rather than left lying around as an invitation to reintroduce the hazard.
- The in-lock re-proof stays as the merged-safety check. The residual — an
  external ref move between re-proof and delete — is documented at the code
  that accepts it.
- The worktree-conflict guard now wraps EVERY branch deletion, temporary
  t3code/* branches included; a user can check one of those out too. git
  enforces it regardless, so the guard's job is the structured refusal naming
  the directory, and git's own refusal is parsed back into the same shape for
  a worktree that appears in the gap.

The driver test now pins the reason rather than the mechanism: the OID is
unchanged by a checkout, porcelain refuses, plumbing with that very SHA
deletes. Both real-worktree settle tests assert git was never asked, so the
guard is load-bearing and not shadowed by git's own protection.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* test(sessions): give post_report's harness the services settle_session added

Rebase fallout that git could not see: #14's new post_report test builds the
sessions handlers, and those now require GitRepositoryLock and the
source-control registry. Textually clean, semantically broken — the stubs get
both, empty, since post_report reaching either would be a bug.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
…y and settle cascade (#12)

Adds list_sessions (a spawning session's children — status, settled/archived
state, hasReport, provider/model, worktreePath, createdAt) and
archive_session (permanently discard a child, reclaiming its worktree by
default). Archiving subsumes settling via a shared stop-first cascade.

Policy (confirmed with the human owner across review): the 8-child spawn
cap counts only active children — unsettled, or settled with a still-live
provider binding — using actual binding evidence from the provider session
directory rather than trusting session.status (an errored session's
binding can still be live and restartable). A separate 32-child retention
cap bounds total non-archived children regardless of settled state.
list_sessions exposes a state filter (active/settled/all, default active)
sharing the same predicate as the cap, so a spawn_limit_reached refusal is
always explainable by listing state: active.

archive_session is genuinely idempotent (retrying against an
already-archived child is a no-op success) and routes both its unsettled
and already-settled paths through the same binding-authoritative
stop-first check, never deleting a worktree out from under a live or
unconfirmed binding.

Reviewed over five rounds on PR #12 (roughcoder), approved at e72c217fa.

Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
#13)

* feat(sessions): retain queued-delivery receipts

Constraint: Queued delivery must stay observable without synchronizing report timing.
Rejected: Holding reports until queued instructions drain | It would change turn-boundary behavior instead of exposing the race.
Confidence: medium
Scope-risk: moderate
Directive: Add bounded receipt retention before relying on this table for long-lived history.
Tested: focused receipt, reactor, handler, projection, and contract tests; npm run typecheck; root vp run test; vp lint
Not-tested: Browser integration (not requested).

* fix(sessions): make queued receipts pull-first

Constraint: Delivery receipts must not create a parent turn per queued message.
Rejected: Per-receipt parent acknowledgements | They amplify queue bursts and do not survive reactor restart.
Confidence: medium
Scope-risk: moderate
Directive: Keep release attribution explicit and recover stale releasing rows before any new release.
Tested: targeted reactor and handler tests; server typecheck; vp lint
Not-tested: Full suite intentionally deferred to merge gate.

* test(sessions): cover queued receipt recovery

Tested: targeted reactor and migration tests

* test(sessions): lock queued receipt lifecycle

Constraint: Receipt attribution must survive restart and never infer a provider turn positionally.
Rejected: In-memory replay guards | durable projection state and explicit message correlation are required.
Confidence: high
Scope-risk: narrow
Directive: Keep releasing rows distinct from queued terminal cancellations.
Tested: Focused lifecycle and receipt test suites.
Not-tested: Final typecheck and lint still running.

* fix(sessions): complete queued receipt wiring

Constraint: Queued delivery receipts must preserve their transitional state across every read model and test runtime.
Rejected: Treating releasing as queued | it can falsely cancel or positionally consume a delivery.
Confidence: high
Scope-risk: narrow
Directive: Add repository test layers whenever a sessions handler gains a durable dependency.
Tested: apps/server typecheck; full monorepo typecheck.
Not-tested: lint pending.

* fix(sessions): derive queued recovery from events

Constraint: Event-sourced read models may not be repaired by direct projection writes.
Rejected: Direct stale-row updates | they diverge from replayed queue state.
Confidence: medium
Scope-risk: moderate
Directive: Keep provider delivery attribution on the running transition that carries a turn id.
Tested: real SQLite receipt test.
Not-tested: full validation pending.

* test(sessions): exercise durable queued receipt storage

Constraint: Receipt queries must run against a real SQLite projection, not only repository stubs.
Rejected: Projection-authored stale recovery | replayed events remain the single source of truth.
Confidence: high
Scope-risk: moderate
Directive: Keep stale releasing recovery event-driven.
Tested: targeted reactor, decider, and real SQLite receipt tests.
Not-tested: full typecheck and lint pending.

* fix(sessions): integrate receipts after main rebase

Constraint: Usage snapshots and report amendments landed while receipt delivery was under review.
Rejected: Dropping either contract during rebase | ping and read must expose both features.
Confidence: high
Scope-risk: narrow
Directive: Receipt reads remain bounded, thread-scoped, and best-effort.
Tested: targeted sessions, reactor, decider, and real SQLite receipt tests.
Not-tested: full typecheck and lint pending.

* fix(sessions): reconcile reactor fixtures after rebase

Constraint: Preserve usage snapshots and queued delivery receipts across the final wave rebase.
Rejected: Omitting usage queries from the reactor fixture | terminal report tests then never reach report dispatch.
Confidence: high
Scope-risk: narrow
Directive: Keep session-reactor fixtures current with required snapshot-query dependencies.
Tested: npm run typecheck; focused contracts, sessions, real-DB receipt, reactor wiring tests; vp check.
Not-tested: Full repository test suite.
…eld (#19)

The receipts feature (#13) added queuedDeliveryMessageId to the session
projection; the two complete-session expectations in the hydration test
were merged without it (each PR's gate passed individually; the combined
fixture gap only surfaced on main).

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
* fix(orchestration): prevent queued delivery replay loops

Persist the delivery marker across session snapshots and bound stale redelivery retries so a failed attribution cannot repeatedly start paid turns.

Constraint: preserve event-sourced queue state and replay compatibility\nRejected: projection-only recovery | it diverges from the command read model\nConfidence: high\nScope-risk: narrow\nDirective: keep queued delivery attribution correlated by message id through every session projection\nTested: npm run typecheck; pnpm exec vp lint; targeted queued delivery tests (188 passed)\nNot-tested: full repository suite

* fix(orchestration): retain valid queued delivery attribution

Prevent unrelated rebinds and recovery timing from erasing or misattributing an in-flight queued delivery marker.

Constraint: preserve message-id correlation without attributing queued rows\nRejected: positional queued-row attribution | it can consume the wrong delivery\nConfidence: high\nScope-risk: narrow\nDirective: only consume a queue row after it is explicitly releasing\nTested: npm run typecheck; pnpm exec vp lint; focused queued attribution tests\nNot-tested: full repository suite
Detect provider runtime death during an active turn and transition through the event-sourced terminal path.\n\nReviewed and validated by Codex through the Phoenix harness.
Ship the new phoenix artwork as the production and local-dev mark so desktop, web, and mobile share one icon instead of the old T3/feather tile.

Co-authored-by: Cursor <cursoragent@cursor.com>
Adds client-local VS Code Remote-SSH editor handoff for remote environments while preserving Phoenix preview and safe fallbacks. Implemented with Codex in Phoenix.
roughcoder and others added 23 commits September 7, 2026 09:34
feat(mobile): match desktop navigation and native action dialogs
* fix(web): correct sidebar surfaces spacing and activity indicators

* fix(web): tighten sidebar search and section spacing

* fix(web): size footer buttons to their active labels
* fix(mobile): restore pull to refresh on session lists

* fix(mobile): handle reconnects and partial session refresh failures

* test(mobile): cover refresh before connection state resolves
* feat(web): build environment management tabs and journeys

* fix(runtime): avoid restarting newly paired environments

* fix(desktop): cover environment appearance in settings persistence

* test(web): tolerate additional settings row attributes
* fix(web): align environment tables with usage density

* feat(web): add search to every environment table

* fix(web): preserve unavailable provider recovery actions

* fix(web): resolve default provider enabled state

* fix(web): guard provider edits and search displayed table values

* test(web): exercise provider restore permission with saved account
* feat(web): build the schedules destination and sidebar

* fix(web): align schedules layout with Paper measurements

* fix(web): refine schedules visual states and recovery actions

* fix(web): finish schedules browser visual review

* fix(web): preserve schedule failure diagnostics in history
* feat(server): prompt agents to suggest session orchestration

* feat(mobile): expose session orchestration settings

* fix(mobile): clarify orchestration setting availability
* fix(web): improve desktop update controls and release notes

* fix(web): clarify downloading release notes and reuse update helpers
Expose the existing environment setting in the shared desktop and web General panel, with search targeting and reset support. Document where to find it.

Validation: 26 targeted settings tests, web typecheck, targeted lint and format checks, and production web build passed. No interactive UI verification performed.
…g a dead reconnect

A paired client's session lasts 30 days and there is no renewal path, so
every remote environment eventually hits expiry. The server knew the
credential had expired, but the wire error flattened that into
`invalid_credential`, and the client rendered a generic red "<label> is
offline" banner whose Reconnect button re-sent the same dead credential
forever. The only visible difference from an ordinary network outage was
the banner's colour.

Expiry now survives the trip. `EnvironmentAuthInvalidError` carries an
optional `credentialExpired` flag, set when the underlying cause is a
session or websocket expiry. It is optional and only sent when true, so
older clients keep decoding the payload and fall back to `reason`.

The clients act on it. `presentConnectionState` exposes `blockedReason`,
and the shared `connectionNeedsPairing` helper answers the question both
surfaces were getting wrong: retrying cannot clear an authentication
block. Web swaps the banner to "<label> needs pairing again" and drops
the Reconnect button; mobile does the same and stops promising it "will
keep retrying automatically".

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
@gabrielsayers1

Copy link
Copy Markdown
Author

Opened against the wrong repository by mistake — this targets the goodbirdhq/phoenix fork, not upstream. Reopening there. Apologies for the noise.

@github-actions github-actions Bot added vouch:unvouched PR author is not yet trusted in the VOUCHED list. size:XXL 1,000+ changed lines (additions + deletions). labels Sep 12, 2026
@macroscopeapp

macroscopeapp Bot commented Sep 12, 2026

Copy link
Copy Markdown
Contributor

Macroscope skipped reviewing this pull request. Per-review cost limit exceeded (workspace setting).

This review would cost an estimated $63.06, which exceeds your per-review limit of $15.00.

The top 3 files driving up this estimate:

File Diff Size Estimate
apps/server/src/mcp/toolkits/sessions/handlers.ts 116.87KB $5.84
docs/internals/session-orchestration.md 55.14KB $2.76
packages/contracts/src/sessionOrchestration.ts 46.26KB $2.31

Tip

To get this pull request reviewed, you can:

  1. Comment @macroscope-app on this PR to request a manual review (monthly spend limits still apply).
  2. Exclude the file(s) above from review by adding a pattern to your .macroscope/ignore.md — note that creating this file replaces Macroscope's built-in default ignores rather than extending them.
  3. Raise your cost limit in your workspace billing settings.

Turn off this reminder going forward

@macroscopeapp

macroscopeapp Bot commented Sep 12, 2026

Copy link
Copy Markdown
Contributor

Approvability

Verdict: Not approved

Macroscope's review found this PR not approvable — The PR is a broad fork/rebrand and capability expansion touching authentication, agent execution, persistence, desktop/mobile behavior, and production release workflows—not just an isolated credential-message fix. It also changes product defaults and suppresses a static-analysis job, so the combined runtime and operational impact requires human review.

Not approved because:

  • Per-review cost limit exceeded (workspace setting). Approvability relies on correctness review in order to determine eligibility

Review your spending limits in Billing settings, or comment @macroscope-app review this PR to bypass the limit and review now. You can add or adjust custom eligibility rules. Learn more.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

size:XXL 1,000+ changed lines (additions + deletions). vouch:unvouched PR author is not yet trusted in the VOUCHED list.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants