Repository navigation
fix(auth): explain expired environment credentials instead of offering a dead reconnect - #11462
gabrielsayers1 wants to merge 241 commits into
Conversation
* feat: agent sessions can spawn and orchestrate other sessions
A running agent session can now create sibling sessions on any configured
provider, exchange messages with them, and receive their completion reports
without polling. New sessions MCP toolkit (list_session_providers,
spawn_session, send_to_session, read_session, stop_session, post_report) on
the per-thread t3-code server; spawned threads get their own worktree by
default and carry a spawnedByThreadId link. A SessionSpawnReactor starts a
turn on the spawning thread when a child posts its report or hits a provider
error. Reports persist as a projection and render as cards; child threads
show a Spawned-by banner. Guardrails: direct-children-only scoping, 8-child /
3-deep caps, child permission mode capped at the parent's, and a Settings
toggle (enableSessionOrchestration) that applies to running sessions
immediately.
Also extracts the thread bootstrap macro from ws.ts into a shared
ThreadTurnBootstrap service, bumps @anthropic-ai/claude-agent-sdk to 0.3.226
(new SDK message types handled in ClaudeAdapter), and fixes an MCP gotcha
where an empty tool parameter struct made Claude Code drop the entire
server's toolset.
Built with Claude Fable 5 on Claude Code.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* feat: spawn_session accepts per-model provider options
list_session_providers now returns each model's option descriptors
(reasoning effort, context size, …) and spawn_session takes an optional
options array of {id, value} selections, validated against the chosen
model's descriptors before being carried into the thread's model selection.
Verified live: a Claude session spawned a Codex child with
reasoningEffort=low; the selection persisted on the spawned thread and the
completion report woke the parent.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
---------
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
Renames the fork's user-visible identity from T3 Code to Phoenix and gives it a distinct OS-level identity, so it installs and runs alongside upstream T3 Code rather than colliding with it. Runtime identity diverges, source identity does not: state dir (~/.phoenix), env var (PHOENIX_HOME), desktop userData, product name, URL scheme (phoenix://), bundle ID (com.goodbird.phoenix), Linux desktop entry and WM class, systemd unit, ports (3873, 13873/5833), CLI binary and MCP server id all change. The @t3tools package scope, t3-prefixed paths and internal symbols stay as upstream has them so merges stay cheap. See docs/internals/branding.md. Notably, the legacy-userData adoption step is removed entirely - keeping it would have pointed two live applications at one directory - and PHOENIX_HOME deliberately has no T3CODE_HOME fallback, unlike every other variable, because the base dir holds the SQLite database and auth state. Also reworks CI for the fork. Upstream targets Blacksmith runners, which a fork without a Blacksmith installation queues indefinitely rather than failing, so no check on this repository had ever executed. Points ci.yml at GitHub-hosted runners, disables the nine workflows that need T3's infrastructure (GitHub-side, so the files stay byte-identical to upstream), and adds phoenix-build.yml to produce an unsigned macOS DMG per merge with a GitHub pre-release and Slack notification. That first real CI run surfaced three defects that had been invisible on main, all from PR #1 landing with zero executed checks: a missing `reports` field in a mobile test fixture, a missing `spawnedByThreadId` in two ProjectionSnapshotQuery assertions, and an upstream image-compression test that starves under parallel load. All three are fixed here. Attribution to T3 Code is added to the README and Settings -> About; the original MIT copyright is untouched.
The first Phoenix Build produced a DMG macOS refuses to open, reporting
"Phoenix (Alpha) is damaged and can't be opened".
Nothing was damaged. Passing CSC_IDENTITY_AUTO_DISCOVERY=false makes
electron-builder skip signing altogether, which leaves the bundle carrying
Electron's stock linker-signed signature - identifier "Electron", no sealed
resources - while the bundle around it has been rewritten. codesign --verify and
spctl both reject that state ("code has no resources but signature indicates
they must be present"), and on Apple silicon macOS refuses to launch it.
Supplies codesign's ad-hoc identity ("-") for unsigned macOS builds instead, and
stops disabling identity auto-discovery on macOS so electron-builder actually
runs the signing step. Ad-hoc signing needs no certificate and no Apple account,
and signs inside-out - helpers and frameworks before the outer app - which a
post-hoc `codesign --deep` does not do correctly.
Also corrects the release notes, which told users to clear the quarantine
attribute. That was the wrong remedy: the fault was an invalid signature, not
quarantine, so xattr would not have helped. The notes now describe the
notarisation prompt an ad-hoc signed build actually produces, and flag "is
damaged" as a bug to report rather than something to work around.
Swaps T3's wordmark for the feather-in-flame artwork across the production channel: macOS/iOS/universal 1024 PNGs, the Windows ICO, and the web favicons. Framing is cropped so the artwork fills 94% of the icon height rather than 81%. The mark is tall and narrow (3080x4332 in a 5333 canvas), so the remaining whitespace is at the sides and is a property of the shape, not the framing - cropping further would leave no breathing room top and bottom. The background stays white. A dark background also works, since the black quill reads against the orange flame rather than against the backdrop, but white is what was chosen for now and is a one-line change in icon.json to revisit. The assets were generated directly rather than through `pnpm icons:export`: that pipeline needs Icon Composer 2.x for design generation 26, and Xcode 26.4.1 ships version 1.4. The .icon project is updated to reference the new artwork so the pipeline stays correct for whoever has a compatible Icon Composer, but the committed PNGs are what the desktop build actually consumes - stageMacIcons feeds the 1024 PNG through sips and iconutil at package time. `icons:check` is not run in CI, so the two will not conflict. Development and nightly channels still carry the T3 artwork, which incidentally makes local dev builds easy to tell apart from a release build in the Dock.
First launch of the packaged app asked to unlock "t3code Safe Storage", not a
Phoenix-owned item. Electron derives app.getName() from the staged app's
package.json name, and macOS safeStorage keys its Keychain entry off that
("<name> Safe Storage"). The staged name was still "t3code", so Phoenix and
upstream T3 Code encrypted their secrets against a single shared keychain item -
the exact class of collision the separate runtime identity exists to prevent.
Sets the staged name to "phoenix". This is distinct from the workspace package
name in apps/server/package.json, which stays "t3" because it never reaches the
OS.
docs/internals/branding.md claimed the Keychain item followed productName and
would therefore be "Phoenix (Alpha) Safe Storage". That was wrong on both counts
and had this identifier filed under source identity, which is why it was missed.
Corrected.
Also fixes the release notes, which told users to Control-click and choose Open.
macOS 15 removed that bypass for un-notarised apps, so System Settings is now
the only route; the notes describe that flow and the actual dialog text.
The desktop sidebar still read "T3 Code" after the rebrand. The name was never a string: SidebarBrand rendered an inline SVG of the "T3" glyph followed by a separate hardcoded "Code" literal, so no search for "T3 Code" could ever have matched it. The mobile home header and compact title did the same. Drives the desktop sidebar from APP_BASE_NAME so it follows the branding constant rather than hardcoded artwork, and uses the name directly on mobile. Deletes T3Wordmark, which has no remaining callers. This is the third rebrand miss found by running the app rather than by reading it, after the broken signature and the shared keychain item. All three were invisible to grep: artwork, an OS-derived identifier, and a split literal.
A running Phoenix install bound port 3774 while 3873 sat free: it had tried 3773, found upstream T3 Code already there, and scanned forward. Adjacent ports look harmless, but the order decides the outcome - start Phoenix first and it takes 3773, pushing T3 Code off its own port. ServerConfig.DEFAULT_PORT was moved to 3873 in the rebrand, but the desktop keeps its own constant and never consulted it, so the change had no effect on the packaged app. packages/ssh had a third copy for remote tunnels. Points DEFAULT_DESKTOP_BACKEND_PORT and DEFAULT_REMOTE_PORT at 3873, and corrects the scan-range comment in the server's auth notes. Found by inspecting a running install rather than by reading the code - the same way the broken signature, the shared keychain item and the T3 wordmark surfaced.
Normalize the feather source for Icon Composer and regenerate the production macOS artwork with a rounded, inset body so Dock rendering no longer exposes an opaque square. Constraint: classic macOS icons require an inset body with transparent canvas around it. Rejected: patch the installed signed app bundle | release assets must come from tracked source. Confidence: high Scope-risk: narrow Directive: preserve the transparent macOS safe area in future production icon exports. Tested: native Icon Composer iOS/macOS renders; iconutil ICNS packaging; extracted 1024px rendition; git diff --check. Not-tested: installed release bundle; no release was requested.
Sync the fork with pingdotgg/t3code through 63e6fae. Notable upstream additions: Open VSX theme search (pingdotgg#5654), OKLCH theme palettes (pingdotgg#6036), hourly past-24-hour usage view (pingdotgg#6170), mobile thread title regeneration (pingdotgg#6253), and Windows/Azure DevOps/GitLab fixes. One conflict, in apps/web/src/components/settings/ThemeEditorPanel.tsx. Upstream pingdotgg#6183 makes environment artwork theme aware and removes the per-theme opt-in toggle; our rebrand had only relabelled that toggle's copy to "Show Phoenix environment artwork". Resolved in favour of upstream, since the surrounding sidebarArtwork state it referenced is gone. The themePalette data model upstream retains is untouched. Verified: pnpm typecheck (no errors), pnpm test (2381 passed, 7 skipped). Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
… documented limits (#9) * feat(sessions): document messageLimit bounds, add structured report fields, fix stop_session null result - read_session/post_report tool descriptions now state messageLimit's 0-20 range, 5 default, and 16,384-char per-message truncation (the schema already enforced min/max; only the description was missing it). - post_report accepts optional findings/validation/recommendation/ completionPercent fields alongside the markdown summary, threaded through the command, event, and a new JSON column (migration 043) so read_session can surface them. All additive, no breaking changes. - stop_session now returns { threadId, status } instead of null, which was failing MCP client schema validation ("expected record, received null") even though the stop succeeded server-side. - spawn_session's description documents the SESSION_SPAWN_MAX_CHILDREN cap, and the spawn-limit error message now correctly says archiving (not stopping) is what frees a slot. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * fix(sessions): bound structured report payload size, degrade gracefully on bad JSON Addresses PR #9 review feedback: - findings/validation.performed/validation.gaps arrays were unbounded while scalar fields were capped, letting a huge findings array store and return wholesale into a parent's context via read_session. Cap findings to 100 entries, performed/gaps to 50 each, and add a 32KB combined encoded-size cap across findings/validation/recommendation/completionPercent as defense in depth. - A malformed structured_json blob previously failed the whole DB row via strict schema decoding, which would have blocked read_session and thread detail hydration. Decode it leniently instead (new decodeStructuredReportFields helper): parse/schema failures log a warning and are treated as absent, never blocking report access. - Services/ProjectionThreadReports.ts's completionPercent was a plain Schema.Int with no 0-100 bound; it now reuses SessionReportStructured's fields directly so persistence-layer bounds can't drift from the contract. - encodeStructured now writes SQL NULL instead of "{}" when there's nothing to store, keeping "no structured data" and "empty structured data" distinguishable in the column. - Renamed handlers.test.ts to tools.test.ts (it only exercised ReadSessionInput/PostReportInput schema decoding, not any handler), and added a focused decodeStructuredReportFields.test.ts plus size/array-cap cases to tools.test.ts. Migration numbering and #8 untouched, per review instructions. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> --------- Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
* feat(sessions): let spawned agents inspect a specified revision Constraint: Keep checkout creation inside the existing bootstrap rollback boundary.\nRejected: Ad-hoc shell commands in the MCP handler | GitWorkflowService already exposes the Git driver seam.\nConfidence: high\nScope-risk: narrow\nDirective: Worktree cleanup and reclamation remain a separate lifecycle concern.\nTested: vp test run packages/contracts/src/sessionOrchestration.test.ts apps/server/src/mcp/toolkits/sessions/handlers.test.ts; npm run typecheck.\nNot-tested: Live provider/session spawn against a remote pull request. * fix(sessions): keep checkout metadata best effort Constraint: Session reports and controls must survive missing or unreadable worktrees.\nRejected: Making checkout metadata a required read/spawn prerequisite | optional fields must not orphan live sessions.\nConfidence: high\nScope-risk: narrow\nDirective: PR-ref portability beyond GitHub remains follow-up work.\nTested: vp test run apps/server/src/mcp/toolkits/sessions/handlers.test.ts apps/server/src/orchestration/ThreadTurnBootstrap.test.ts packages/contracts/src/sessionOrchestration.test.ts; npm run typecheck.\nNot-tested: Live remote fetch and provider-session spawn. * test(sessions): type checkout failure doubles accurately Constraint: Tests must match nominal Git workflow error and status contracts.\nRejected: Untyped Error and partial result stubs | root tsgo rejects the incompatible channels.\nConfidence: high\nScope-risk: narrow\nDirective: Keep optional checkout enrichment non-blocking.\nTested: npm run typecheck.\nNot-tested: Focused Vitest suite not rerun; this change is test typing and assertions only. * fix(sessions): use typed checkout failure channels Constraint: Effect code must use platform services and tagged failure values.\nRejected: Node path import and generic Error failures | root tsgo treats them as errors.\nConfidence: high\nScope-risk: narrow\nDirective: Preserve the fetch-versus-missing-ref distinction through typed errors.\nTested: npm run typecheck; vp run --last-details (all 15 tasks succeeded, exit 0).\nNot-tested: Focused Vitest suite not rerun. * test(server): resolve bootstrap revisions in seam harness Constraint: Bootstrap now resolves a commit before creating a worktree.\nRejected: Leaving the driver mock implicit | seam tests must exercise the production dependency path.\nConfidence: high\nScope-risk: narrow\nDirective: Keep revision stubs transparent so base-ref assertions remain meaningful.\nTested: vp test run apps/server/src/server.test.ts -t five bootstrap seam cases; vp test run apps/server/src/mcp/toolkits/sessions/handlers.test.ts apps/server/src/orchestration/ThreadTurnBootstrap.test.ts packages/contracts/src/sessionOrchestration.test.ts; npm run typecheck (exit 0).\nNot-tested: Full t3 server suite.
…#6) * feat: let spawned sessions stop gracefully Constraint: Stop requests must remain command-to-event-to-projection flows.\nRejected: A parallel stop-state store | it would bypass durable session projection.\nConfidence: medium\nScope-risk: moderate\nDirective: Keep MCP tool input schemas explicitly typed; no empty structs.\nTested: targeted immediate-stop reactor test; focused migration, contract, and decider tests.\nNot-tested: full typecheck remains blocked by the machine's Node/dependency setup and the resource cap. * fix: preserve graceful stop invariants Persist grace deadline episodes so restart recovery and a resumed session cannot be stopped by stale work. Constraint: graceful stops must survive reactor restarts without blocking ordinary session resume. Rejected: in-memory daemon-only deadline | lost on restart and could stop a new session. Confidence: high Scope-risk: moderate Directive: keep grace notices explicitly tagged; do not gate normal turn starts on stopped status. Tested: focused graceful deadline and stopped-session resume tests; decider, sessions handler, and migration tests; npm run typecheck (exit 0). Not-tested: provider-specific steer behavior. * fix: keep rebased graceful stop typed Constraint: session stop must coexist with the merged report and checkout contracts. Rejected: retaining pre-rebase type assumptions | stale contract and Effect diagnostics failed typecheck. Confidence: high Scope-risk: narrow Directive: keep stop-session schema fixtures branded and time through Effect DateTime. Tested: npm run typecheck -- --pretty false; targeted session, decider, reactor, and migration tests. Not-tested: provider-specific steer behavior. * test: cover graceful stop session defaults Constraint: projected session snapshots include stop-audit and graceful-deadline defaults. Rejected: partial fixture update | snapshot hydration compares the complete session shape. Confidence: high Scope-risk: narrow Directive: update complete session fixtures when projection fields are added. Tested: targeted ProjectionSnapshotQuery, sessions, decider, reactor, and migration tests (95 passed). Not-tested: full typecheck; prior gate run exited 0.
#5) * feat(server): make session messages queue safely Persist busy-child deliveries as FIFO queued turn starts and release them at provider session boundaries. Interrupt mode requests the existing provider interrupt before the queued replacement is released. Constraint: send_to_session must remain event-sourced and provider-neutral. Rejected: provider-level notify steering | adapter semantics are inconsistent and need an explicit capability contract. Confidence: high Scope-risk: moderate Directive: add notify only after every provider has an explicit steering or fallback contract. Tested: pnpm exec vp test run apps/server/src/mcp/toolkits/sessions/handlers.test.ts apps/server/src/orchestration/decider.settled.test.ts (15 passed); npm run typecheck under Node 24 found one startup error-channel leak, fixed afterward. Not-tested: full typecheck was not rerun due the session's one-run resource constraint; no browser or full-suite validation. * fix(server): keep queued session delivery live Make queued delivery release idempotent, recover stale persisted sessions only after checking live provider bindings, cancel terminal queues, and bound interrupt fallback with parent-visible cancellation. Constraint: queued delivery must remain event-sourced and safe across restart, provider failure, and concurrent release signals. Rejected: unconditional stop-then-send fallback | provider exit can race terminal queue cancellation and resurrect stopped work. Confidence: high Scope-risk: moderate Directive: preserve queued IDs in the command read model when changing turn projections. Tested: pnpm exec vp test run apps/server/src/orchestration/Layers/SessionSpawnReactor.test.ts apps/server/src/orchestration/decider.settled.test.ts apps/server/src/mcp/toolkits/sessions/handlers.test.ts apps/server/src/orchestration/projector.test.ts apps/server/src/orchestration/Layers/ProjectionSnapshotQuery.test.ts (50 passed); npm run typecheck under Node 24 found four localized typing issues, fixed afterward. Not-tested: typecheck was not rerun due the one-run resource constraint; no full suite or browser validation. * fix(server): recheck queued session liveness Move provider liveness observation into the serialized recovery decision and make post-dispatch acknowledgement uncertainty explicit. Constraint: Shared-machine validation permits one typecheck run and targeted tests only. Rejected: Two-observation failure detection | a decision-time in-memory provider check closes the stale snapshot window directly. Confidence: high Scope-risk: narrow Directive: Keep provider liveness checks inside serialized recovery decisions. Tested: pnpm exec vp test run apps/server/src/orchestration/Layers/SessionSpawnReactor.test.ts apps/server/src/mcp/toolkits/sessions/handlers.test.ts apps/server/src/orchestration/decider.settled.test.ts (21 passed, exit 0). Not-tested: npm run typecheck was not rerun after fixing its two exact diagnostics; its single permitted run exited 1. * fix(server): document recovery binding invariant Record the provider-directory assumption at the recovery decision and keep recovery sweep interruption out of the error channel. Constraint: Recovery may force-release only when directory absence is authoritative. Rejected: Optional cleanup refactors | closing the reviewed invariant and compiler issue keeps this commit narrow. Confidence: high Scope-risk: narrow Directive: Revisit the recovery failure detector if provider directory absence can become transient during rebind. Tested: Node 24 npm run typecheck (exit 0); targeted SessionSpawnReactor test (6 passed, exit 0); git diff --check. Not-tested: Full test suite and browser validation were not run. * fix(server): preserve graceful stop steering Keep grace-stop notices immediate while ordinary busy-session messages queue, and make provider reactor fixtures model a completed turn before follow-up commands. Constraint: Rebase must preserve graceful-stop deadlines and queued session delivery semantics. Rejected: Queueing grace notices | it can hide the stop deadline until the turn it must stop has already ended. Confidence: high Scope-risk: narrow Directive: Internal in-turn steering commands must opt out explicitly if future delivery defaults queue busy sessions. Tested: Node 24 npm run typecheck (exit 0); targeted provider/handlers tests (53 passed, exit 0); targeted decider/session reactor/snapshot tests (39 passed, exit 0); git diff --check. Not-tested: Full repository test suite and browser validation were not run. * test(server): preserve pending steer coverage Exercise conflicting provider turn acceptance through the intentional immediate grace-notice steer now that ordinary busy-session messages queue by default. Constraint: The ingestion test must create a pending turn start before asserting conflicting turn acceptance. Rejected: Restoring implicit busy-session steering | it would violate send_to_session's queue-by-default contract. Confidence: high Scope-risk: narrow Directive: Pending-turn ingestion tests must use an explicit immediate-start path when the thread is already running. Tested: Node 24 npm run typecheck (exit 0); targeted ProviderRuntimeIngestion, ProviderCommandReactor, SessionSpawnReactor, and decider suites (112 passed, exit 0); single regression test (exit 0); git diff --check. Not-tested: Full repository suite and browser validation were not rerun locally.
…e_session with worktree cleanup (#10) * feat(orchestration): synthesize terminal reports and add settle_session A spawned session that died without calling post_report left its parent with nothing to act on: a `stopped` child produced no notification at all, and an errored one produced a bare error message with no record of what it had done. Separately, nothing in the server ever reclaimed a spawned worktree, so a long orchestration run leaked a directory per child. Synthetic terminal reports: when a spawned child's session reaches `stopped` or `error` with no report posted, SessionSpawnReactor dispatches an ordinary `thread.report.post` carrying the termination reason, the last tool activity and assistant message, and an explicit "work is likely unfinished" warning. Everything downstream — projection, the report card, the parent wake-up — behaves as it does for an agent-posted report, which is how `stopped` now notifies the parent at all. `SessionReport.origin` ("agent" | "system", migration 043) keeps a synthesized report from ever reading as the child's own claim, in the parent's notification and in the UI badge. settle_session: the parent explicitly settles a child rather than sessions settling themselves. It refuses a starting/running child with an actionable message instead of stopping it, and keeps the decider's guards for open approvals and queued turns. `cleanupWorktree: true` permanently deletes the child's worktree and its temporary `t3code/…` branch — refused, with the specific dirty files and unpushed commit count, unless the work is committed and pushed or `force: true` is passed. Branches Phoenix did not create are kept and reported. The result always names what was removed and what was kept, and read_session now exposes the child's worktree path so a deleted worktree stops being advertised. Validation: `npm run typecheck` clean; new handler and reactor tests plus the decider report tests pass. * fix(orchestration): stop a live session on settle, mark reports only once posted Addresses code review on PR #10. [major] The terminal-report dedup set was marked before the dispatch that posts the report. processEventSafely swallows dispatch failures, and a terminated session emits no further status transition, so a transient failure branded the episode "handled" with nothing persisted and the parent was never told — worse than the old behaviour, which at least never pretended. The episode is now marked only once the report is actually posted, the persisted `reports` check remains the real duplicate guard, and the dispatch retries twice before giving up, since this is the last chance to reach the parent. [major] settle_session now stops a still-alive session instead of leaving a `ready` provider process running behind a settled thread. A turn in flight (starting/running) is still refused — interrupting live work stays a deliberate stop_session — but an idle-but-alive child is stopped as part of settling, because settling is the parent declaring it finished. Order is stop -> settle -> assess -> remove, which also closes the window the reviewer flagged: the dirty/unpushed check now runs after the process is gone, so nothing can write to the worktree between the check and the removal. If the session does not reach "stopped" within the timeout the thread is still settled but cleanup is withheld and the caller told. [minor] Documented that deleteRef is a trusted-caller primitive: the branch-safety decision lives in decideBranchCleanup, not in the driver. Tests: both majors were in wiring that only pure-helper tests covered. Added SessionSpawnReactor.wiring.test.ts driving the real reactor against a stub engine (verified failing against the old ordering: 3 dispatch attempts instead of 6), and settleSession.test.ts driving the real handler against stub services, asserting the stop/settle/inspect/delete call order. Validation: `npm run typecheck` exit code 0; 34 tests across the five touched files pass; `vp lint` clean. * fix(contracts): supply required stop-audit fields in the checkout result test Pre-existing on main, surfaced by rebasing onto it: #4's ReadSessionResult decode fixture does not supply the stoppedBy / stopRequestedAt / stopReason / interruptedToolCall / lastCompletedOperation fields that #6 made required, so `packages/contracts/src/sessionOrchestration.test.ts` fails at origin/main with "SchemaError: Missing key at [stoppedBy]". Verified by running the test against main's own unmodified sessionOrchestration.ts. Not caused by this branch, but it would land on top of it, so it is fixed here rather than left for the merge.
… sibling access (#8) Rebased onto main after #4/#5/#6/#9/#10 merged; carries the reconciliation: - Envelope delivery applies to agent AND system-origin synthetic reports (same report-posted path); formatReportMessage keeps #10's origin-aware lead and delivers >1KB summaries as an envelope with a read_report hint. - SessionReportEnvelope carries #9's structured data compactly: recommendation and completionPercent whole, findings/validation as counts; and #10's origin so a synthesized epitaph is visible at a glance. read_report returns the full findings/validation arrays with every page (bounded by the 32KB structured cap) plus origin. - findByReportId now selects abstract, structured_json, and origin and maps through the shared mapReportRow, so no column can be silently dropped by one read path; the ProjectionSnapshotQuery Struct.pick allowlist gained abstract. - Migration renumbered 043 -> 046 (043 structured, 044 stop audit, 045 origin landed first). - read_report docs record UTF-16 code-unit paging (pages can run one unit short or long at surrogate boundaries) and the unguessable-UUID timing assumption behind the reportId-only lookup path. Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
Expose the in-memory background-liveness and plan-progress state (already used by the sidebar) over MCP without starting a turn or disturbing the child: a new ping_session tool returns session status, settlement, last provider activity timestamp, current background activity, plan progress, whether a report has landed, and a snippet of the last assistant message. read_session gains the same lastActivityAt/currentActivity fields (additive, optional) for parity. hasReport and lastAssistantMessage are served by two purpose-built, bounded ProjectionSnapshotQuery reads (getThreadHasReport, getLastAssistantMessage) rather than a full/windowed thread detail snapshot, so ping_session stays cheap regardless of a child's history, and the last-assistant-message lookup is correct even when the child's newest turn happens to be user-only. All optional enrichment (provider session directory lookup, report existence, last message) is best-effort: a lookup failure degrades to null/false rather than failing the call. Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
…h merged session features (#11) Assert graceful-stop audit defaults explicitly and run Effect-based tests through @effect/vitest so main's release gate matches the merged runtime behavior. Constraint: Restore main CI with one narrow test-only PR. Rejected: Removing stop audit metadata | the client command intentionally propagates user stop attribution. Confidence: high Scope-risk: narrow Directive: Use @effect/vitest for tests that execute Effect values; do not call manual runtimes. Tested: vp lint (exit 0); Node 24 npm run typecheck (exit 0); focused four-file tests (102 passed, exit 0); root vp run test after ensuring Electron runtime (2485 passed, 7 skipped, exit 0); git diff --check. Not-tested: Browser and mobile device integration were not run.
* feat(contracts): add sidebarSessionHierarchyEnabled client setting A rendering mode over the spawnedByThreadId links the server already stores, so it lives with the other sidebar client settings rather than in ServerSettings: nothing about it is server-authoritative. Defaults off. Turning it on changes how every existing thread list is laid out, so it opts in rather than surprising users on upgrade. The desktop client-settings fixture builds a complete ClientSettings literal, so every new key has to be added there too. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * feat(web): build the sidebar session tree from spawnedByThreadId buildSidebarThreadHierarchy nests threads under the session that spawned them, recursively and to any depth. Sibling order is the incoming order: the caller has already applied the section's sort, and hierarchy mode re-parents rows without reordering them. Two invariants matter more than the nesting itself, because both failure modes hide work from the user: - Every input thread is emitted exactly once. A parent that isn't in the list (archived, snoozed, settled into another section, filtered out by project scope) leaves its child at top level instead of dropping it, and threads caught in a parent cycle are emitted flat for the same reason. - Parent links are matched per environment. The link is a bare thread id, so two environments can hold the same id without being related. Indentation caps at four levels while depth keeps counting, so a deep tree stops stealing horizontal room from titles but still reads as nested. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * feat(web): render nested sessions as indented, compact rows A spawned thread renders under its parent, indented one step per level with a hairline rail down the gutter so the nesting is followable without counting pixels. Nested rows drop the branch line and shrink from 4.875rem to 3.375rem. Children still have their own worktrees — the branch goes for density, because a tree of full-height cards buries the parent it belongs to. The signal that line carried moves rather than disappearing: the PR number and terminal indicator ride up beside the status, since the PR is usually the whole point of a spawned session. Hierarchy applies to the inbox only. The other blocks each carry an ordering that nesting would fight: pinned rows are a hand-dragged arrangement, the snoozed shelf sorts by what wakes next, and the settled tail is paginated history where a parent and its children could straddle a page boundary. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * feat(web): add the sidebar hierarchy toggle under session orchestration The switch renders only while session orchestration is on: with it off nothing spawns children, so the mode would have nothing to nest and the row would be a dead end. The gate is on the settings row, not on rendering. An existing tree keeps nesting if orchestration is later turned off, because the nesting is history rather than a live permission. Registered with the settings search index and the dirty/reset-all lists alongside the other client settings. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * docs: document sidebar session hierarchy Covers what nests, the four-level indent cap, why child rows are shorter (and that children keep their own worktrees), which sections hierarchy applies to, and that a child outlives a parent that leaves the inbox. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
… 047) (#18) * feat(persistence): reserve queued delivery receipts Constraint: Migration 047 must land before migration 048 in the merge train. Rejected: Bundling schema reservation with delivery behavior | migration ordering blocks independent PRs. Confidence: high Scope-risk: narrow Directive: Keep migration 047 limited to durable receipt columns and its query index. Tested: migration test pending. Not-tested: full suite. * fix(persistence): index receipt ordering Constraint: Migration 047's receipt index must serve the bounded query order. Rejected: State-first ordering | it cannot support the receipt timestamp expression. Confidence: high Scope-risk: narrow Tested: migration test pending.
…15) * feat(contracts): add SessionUsageSnapshot for session budgeting Best-effort token/turn accounting so an orchestrating parent can budget instead of flying blind: threaded through PingSessionResult, SessionReport, SessionReportEnvelope, and ReadReportResult. Individual fields are optional since a provider may not report tokens; deliberately no cost estimate, as price tables go stale and tokens are the stable currency. * feat(server): capture usage snapshots in ping_session and posted reports Reuses the existing context-window usage projection (the same projection_thread_activities rows the web client's context-window meter reads) plus two bounded, single-purpose queries (getLatestUsageActivity, getThreadTurnCount) rather than adding new provider plumbing. Usage is computed server-side at ping/report-post time and stored alongside the existing structured report fields, so no migration is needed. Both bounded reads degrade to "field omitted" on failure, matching ping_session's existing enrichment contract. * docs(session-orchestration): document the usage snapshot Explains where the usage numbers come from (the existing context-window projection, not new provider plumbing), why no cost estimate is included, and the report-post-time capture semantics. * fix(sessions): make usage snapshot fields honest, fix turnCount, harden clocks Address REQUEST-CHANGES review feedback on PR #15: - Rename inputTokens/outputTokens to lastTurnInputTokens/lastTurnOutputTokens everywhere: both provider adapters report only the most recent turn's counts at this level, never a session accumulation, and the old names read as spend across the whole session. totalTokens is now sourced only from a provider's own cumulative counter (totalProcessedTokens) and omitted rather than backfilled from context-window occupancy (usedTokens), which is bounded by the window and drops after compaction — the opposite of a monotonic spend number. - getThreadTurnCount now filters turn_id IS NOT NULL, excluding pending turn-start placeholders and queued/interrupting rows that share the table but are not turns. - Guard Date.parse failures in elapsedMsSince/lastTurnDurationMsFromTurn with Number.isFinite so a bad timestamp degrades to 0/null instead of NaN, which would otherwise void the entire structured_json blob (findings/validation/recommendation/completionPercent alongside usage) through the lenient decoder, or fail ping_session's schema encode. - Document cross-provider incomparability of the last-turn token fields (Claude folds cache tokens into input; Codex reports them separately).
…hains (#14) * feat(contracts): optional report amendment links on reports and envelopes post_report gains supersedesReportId so a session can amend an account it already gave — the case that motivated this is a queued instruction landing after the child reported, leaving a stale report as the record. Both fields are optional on SessionReport and on the thread.report.post command / thread.report-posted payload, so every already-persisted report event replays unchanged. supersededByReportId is deliberately absent from the event: at post time no superseding report exists yet, so it is derived on read paths instead. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * feat(server): persist report amendments and derive the reverse link Migration 048 adds projection_thread_reports.supersedes_report_id plus the index every read path uses to resolve the other direction. Only the forward link is stored: an amendment never rewrites the row it supersedes, so the projection stays append-only and "who superseded me" is a correlated subquery rather than a flag that could drift out of sync. Migration id 48 leaves a gap at 47, which a sibling wave-2 slice reserves; ids only have to be ordered, not contiguous. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * feat(sessions): amend reports via post_report supersedesReportId post_report validates the reference before dispatching: it must name a report on the calling thread. A report is a session's account of its own work, so amending another session's report is refused — with the same message as an unknown id, so the denial cannot double as a probe for report ids elsewhere. A dangling link would be worse than a rejection: no reader could follow it. Both ends travel outward. An amending report's parent notification leads with "AMENDED report (supersedes ...)" before the summary, because a parent that already acted on the superseded report has to see that first. read_report on a superseded report still serves its body, plus supersededByReportId and a supersededNotice sentence — a caller paging an old report cannot be relied on to notice a field it was not looking for. The spawned-session instructions now state the rule directly: an instruction arriving after you reported means post an amending report, never claim retroactive compliance. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * docs(internals): record how report amendment works Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix(sessions): keep amendment chains linear, enforced in the decider Review found A→B and A→C was possible, leaving reverse navigation and latest-report selection free to disagree about which report is current. Superseding an already-superseded report is now refused and the caller is handed the head of the chain, which also settles the concurrent-fork race: the second writer loses with an error naming where to re-attach. The authoritative check lives in the decider, not the toolkit. Handler-level validation cannot prevent a fork — two amendments can both pass their pre-checks before either dispatches — whereas the decider runs against the folded read model serialized with command processing, and also covers internally dispatched report posts that never reach the toolkit. The toolkit keeps a pre-check purely for error quality, and re-reads the chain when a dispatch is rejected so a race loser gets the same structured error rather than a generic dispatch failure. Both sides share one implementation and one wording so they cannot disagree. Tests: fork rejection including the race shape, chain-head reporting through a longer chain, a one-turn reactor run asserting AMENDED delivery reaches the parent (not just the formatter), and synthesized-report resurrection — a Phoenix terminal report superseded by the agent's own account after resume. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
…-proven branch deletion (#17) * feat(git): add a per-repository lock for git worktree mutations Git allows one writer per repository: concurrent `git worktree remove` runs on one repo contend for .git/index.lock, and the losers block until the caller's own timeout kills them rather than queueing. This is one Effect Semaphore per repository root so callers can serialize the mutations that take that lock. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * feat(contracts): branch cleanup, settle warnings, and git-hygiene errors settle_session gains `cleanupBranch` (delete a branch Phoenix did not create, against a merge proof), a `branchProof` on the worktree outcome, and a `warning` on the result for a settle whose child process outlived its stop. Two structured errors carry what prose cannot: which leg of the merge proof failed with the SHAs that disagree, and which git lock file is in the way plus the remedy. `ChangeRequest.headRefOid` is the merged commit the proof compares against; optional, since not every host reports one. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * feat(sessions): serialize worktree cleanup and prove branch merges before deleting Settling eight children at once with cleanupWorktree made every cleanup fail: concurrent `git worktree remove` calls fought over .git/index.lock until each timed out, and the one that ran alone was the only one that worked. A timed-out removal then left a zero-byte index.lock that blocked every later git command on the repository. - Worktree removal and branch deletion now run inside GitRepositoryLock, keyed by the project root, so parallel settles queue instead of racing. Sessions still settle independently of cleanup. - A git failure naming a lock file answers with SessionOrchestrationGitLockError: the path, its age, whether it matches the conservative stale heuristic (empty and older than 60s), and the remedy. Phoenix never deletes the lock itself — nothing in this process can prove no live git owns it. - cleanupBranch deletes a user's branch only when local head == remote head == the head commit of a merged PR. `git branch --merged` is useless here: this repo squash-merges, so a merged branch is never an ancestor of main. The proof runs before anything is destroyed, so a refusal costs nothing. - A settle whose stop wait times out no longer succeeds silently; it names the provider and the status the session was last seen in. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * docs(sessions): worktree cleanup git hygiene What changed and why: one cleanup per repository at a time, a leftover lock reported rather than removed (with the heuristic and why Phoenix will not delete it), the squash-merge trap that makes `git branch --merged` the wrong proof, and the settle warning for a process that outlived its stop. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix(sessions): make the lock error reachable, the lock key canonical, and the merge proof race-free Review findings on #17, all four of which were the difference between code that looks right and code that works. - The structured lock error could never fire in production: the driver threw away git's stderr and substituted a fixed detail string, so only a test's injected message ever matched. GitCommandError now carries a bounded stderrExcerpt (tail-first — the fatal line is last), and the driver attaches it to every non-zero exit. Proven end to end against a real held ref lock: real repo → real git failure → structured error naming the real lock file. - The semaphore was keyed on a trimmed path, so a symlinked alias or a linked worktree of the same repository got a *different* permit and serialized nothing. The key is now realpath + the repository's common git directory, with tests for both aliases (they fail against the old keying). - The merge proof was computed before the semaphore and acted on inside it, so a branch updated while a cleanup queued would be deleted on a stale proof. The two heads are re-read inside the critical section immediately before the delete; a branch that moved is kept and reported instead. - Lock matching keys on the .lock path rather than on git's English, and a branch delete that fails on a lock now raises the same structured error instead of swallowing it into a detail string — after clearing the thread's worktreePath, since the directory really is gone by then. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix(vcs): redact credentials from git error excerpts, and delete proven branches by compare-and-swap Round-3 review on #17. - The stderr excerpt I added last round crossed RPC unredacted, and git echoes remote URLs that can carry their own credentials (https://x-access-token:ghs_...@host/o/r). Userinfo, query strings, fragments, and known token shapes are stripped before the excerpt is attached — and before it is truncated, so nothing survives on a boundary. Scheme, host, and path stay: which remote failed is the diagnostic value. This restores an invariant an existing driver test already asserted for args and raw stderr, and which the excerpt had quietly reversed. - The in-lock re-proof is only atomic against writers in this process; a terminal or a second Phoenix can still move the ref between the check and the delete. deleteRef now takes an expectedSha and routes through `git update-ref -d <ref> <old>`, so git itself refuses a moved ref under its own lock. Temporary t3code branches keep the plain force delete: no proof, nothing to hold git to. - A cleanup that removes the worktree but keeps the branch now reports a structured worktree.branchRefusal (branch, reason, SHAs, the expected commit) sharing its reason vocabulary with the BranchNotMerged error, instead of folding a partial success into a prose detail. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix(sessions): refuse to delete a branch another worktree still has checked out Round-4 review on #17. The compare-and-swap closed the ref race but opened a hole: `git update-ref -d` is plumbing and, unlike `git branch -d`, deletes a branch another linked worktree is sitting on — leaving that worktree's HEAD pointing at nothing. Removing our own worktree first is not enough, because the other one can be a worktree a user made by hand that this process has never heard of and the repository lock does not cover. So inside the critical section, before the CAS, the worktree list is consulted and a branch still held anywhere is kept with a structured refusal naming the conflicting path. An unreadable worktree list fails closed for the same reason. The CAS stays as the final atomicity layer. The driver test proves the hazard rather than assuming it: `git branch -d` refuses a checked-out branch, plumbing deletes it. The settle test is the real regression the reviewer asked for — a real repository with a real second worktree, refused; the worktree removed; deleted. Also corrects a comment that claimed argv never carries credentials: clone, fetch, and push take the remote URL as an argument, which is exactly why redaction is written against URL shapes rather than a command allowlist. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix(vcs): delete branches with porcelain, not compare-and-swap Design reversal, and the reasoning I gave last round was wrong. I argued the compare-and-swap bounded the worktree TOCTOU. It does not bound it at all: a checkout does not move the branch's OID, so `update-ref -d <ref> <sha>` succeeds while another worktree holds the branch and leaves that worktree's HEAD dangling. The two remaining races are not equally survivable, and that decides the mechanism. A checkout we get wrong corrupts someone else's directory and no reflog repairs it. A ref move we miss costs a ref pointer the reflog still holds. `git branch -D` refuses a checked-out branch atomically at delete time and cannot see a ref move; the compare-and-swap is exactly the reverse. So: - deleteRef is porcelain again, and the expectedSha parameter is removed rather than left lying around as an invitation to reintroduce the hazard. - The in-lock re-proof stays as the merged-safety check. The residual — an external ref move between re-proof and delete — is documented at the code that accepts it. - The worktree-conflict guard now wraps EVERY branch deletion, temporary t3code/* branches included; a user can check one of those out too. git enforces it regardless, so the guard's job is the structured refusal naming the directory, and git's own refusal is parsed back into the same shape for a worktree that appears in the gap. The driver test now pins the reason rather than the mechanism: the OID is unchanged by a checkout, porcelain refuses, plumbing with that very SHA deletes. Both real-worktree settle tests assert git was never asked, so the guard is load-bearing and not shadowed by git's own protection. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * test(sessions): give post_report's harness the services settle_session added Rebase fallout that git could not see: #14's new post_report test builds the sessions handlers, and those now require GitRepositoryLock and the source-control registry. Textually clean, semantically broken — the stubs get both, empty, since post_report reaching either would be a bug. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
…y and settle cascade (#12) Adds list_sessions (a spawning session's children — status, settled/archived state, hasReport, provider/model, worktreePath, createdAt) and archive_session (permanently discard a child, reclaiming its worktree by default). Archiving subsumes settling via a shared stop-first cascade. Policy (confirmed with the human owner across review): the 8-child spawn cap counts only active children — unsettled, or settled with a still-live provider binding — using actual binding evidence from the provider session directory rather than trusting session.status (an errored session's binding can still be live and restartable). A separate 32-child retention cap bounds total non-archived children regardless of settled state. list_sessions exposes a state filter (active/settled/all, default active) sharing the same predicate as the cap, so a spawn_limit_reached refusal is always explainable by listing state: active. archive_session is genuinely idempotent (retrying against an already-archived child is a no-op success) and routes both its unsettled and already-settled paths through the same binding-authoritative stop-first check, never deleting a worktree out from under a live or unconfirmed binding. Reviewed over five rounds on PR #12 (roughcoder), approved at e72c217fa. Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
#13) * feat(sessions): retain queued-delivery receipts Constraint: Queued delivery must stay observable without synchronizing report timing. Rejected: Holding reports until queued instructions drain | It would change turn-boundary behavior instead of exposing the race. Confidence: medium Scope-risk: moderate Directive: Add bounded receipt retention before relying on this table for long-lived history. Tested: focused receipt, reactor, handler, projection, and contract tests; npm run typecheck; root vp run test; vp lint Not-tested: Browser integration (not requested). * fix(sessions): make queued receipts pull-first Constraint: Delivery receipts must not create a parent turn per queued message. Rejected: Per-receipt parent acknowledgements | They amplify queue bursts and do not survive reactor restart. Confidence: medium Scope-risk: moderate Directive: Keep release attribution explicit and recover stale releasing rows before any new release. Tested: targeted reactor and handler tests; server typecheck; vp lint Not-tested: Full suite intentionally deferred to merge gate. * test(sessions): cover queued receipt recovery Tested: targeted reactor and migration tests * test(sessions): lock queued receipt lifecycle Constraint: Receipt attribution must survive restart and never infer a provider turn positionally. Rejected: In-memory replay guards | durable projection state and explicit message correlation are required. Confidence: high Scope-risk: narrow Directive: Keep releasing rows distinct from queued terminal cancellations. Tested: Focused lifecycle and receipt test suites. Not-tested: Final typecheck and lint still running. * fix(sessions): complete queued receipt wiring Constraint: Queued delivery receipts must preserve their transitional state across every read model and test runtime. Rejected: Treating releasing as queued | it can falsely cancel or positionally consume a delivery. Confidence: high Scope-risk: narrow Directive: Add repository test layers whenever a sessions handler gains a durable dependency. Tested: apps/server typecheck; full monorepo typecheck. Not-tested: lint pending. * fix(sessions): derive queued recovery from events Constraint: Event-sourced read models may not be repaired by direct projection writes. Rejected: Direct stale-row updates | they diverge from replayed queue state. Confidence: medium Scope-risk: moderate Directive: Keep provider delivery attribution on the running transition that carries a turn id. Tested: real SQLite receipt test. Not-tested: full validation pending. * test(sessions): exercise durable queued receipt storage Constraint: Receipt queries must run against a real SQLite projection, not only repository stubs. Rejected: Projection-authored stale recovery | replayed events remain the single source of truth. Confidence: high Scope-risk: moderate Directive: Keep stale releasing recovery event-driven. Tested: targeted reactor, decider, and real SQLite receipt tests. Not-tested: full typecheck and lint pending. * fix(sessions): integrate receipts after main rebase Constraint: Usage snapshots and report amendments landed while receipt delivery was under review. Rejected: Dropping either contract during rebase | ping and read must expose both features. Confidence: high Scope-risk: narrow Directive: Receipt reads remain bounded, thread-scoped, and best-effort. Tested: targeted sessions, reactor, decider, and real SQLite receipt tests. Not-tested: full typecheck and lint pending. * fix(sessions): reconcile reactor fixtures after rebase Constraint: Preserve usage snapshots and queued delivery receipts across the final wave rebase. Rejected: Omitting usage queries from the reactor fixture | terminal report tests then never reach report dispatch. Confidence: high Scope-risk: narrow Directive: Keep session-reactor fixtures current with required snapshot-query dependencies. Tested: npm run typecheck; focused contracts, sessions, real-DB receipt, reactor wiring tests; vp check. Not-tested: Full repository test suite.
…eld (#19) The receipts feature (#13) added queuedDeliveryMessageId to the session projection; the two complete-session expectations in the hydration test were merged without it (each PR's gate passed individually; the combined fixture gap only surfaced on main). Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
* fix(orchestration): prevent queued delivery replay loops Persist the delivery marker across session snapshots and bound stale redelivery retries so a failed attribution cannot repeatedly start paid turns. Constraint: preserve event-sourced queue state and replay compatibility\nRejected: projection-only recovery | it diverges from the command read model\nConfidence: high\nScope-risk: narrow\nDirective: keep queued delivery attribution correlated by message id through every session projection\nTested: npm run typecheck; pnpm exec vp lint; targeted queued delivery tests (188 passed)\nNot-tested: full repository suite * fix(orchestration): retain valid queued delivery attribution Prevent unrelated rebinds and recovery timing from erasing or misattributing an in-flight queued delivery marker. Constraint: preserve message-id correlation without attributing queued rows\nRejected: positional queued-row attribution | it can consume the wrong delivery\nConfidence: high\nScope-risk: narrow\nDirective: only consume a queue row after it is explicitly releasing\nTested: npm run typecheck; pnpm exec vp lint; focused queued attribution tests\nNot-tested: full repository suite
Detect provider runtime death during an active turn and transition through the event-sourced terminal path.\n\nReviewed and validated by Codex through the Phoenix harness.
Ship the new phoenix artwork as the production and local-dev mark so desktop, web, and mobile share one icon instead of the old T3/feather tile. Co-authored-by: Cursor <cursoragent@cursor.com>
Adds client-local VS Code Remote-SSH editor handoff for remote environments while preserving Phoenix preview and safe fallbacks. Implemented with Codex in Phoenix.
feat(mobile): match desktop navigation and native action dialogs
* fix(web): correct sidebar surfaces spacing and activity indicators * fix(web): tighten sidebar search and section spacing * fix(web): size footer buttons to their active labels
* fix(mobile): restore pull to refresh on session lists * fix(mobile): handle reconnects and partial session refresh failures * test(mobile): cover refresh before connection state resolves
* feat(web): build environment management tabs and journeys * fix(runtime): avoid restarting newly paired environments * fix(desktop): cover environment appearance in settings persistence * test(web): tolerate additional settings row attributes
* fix(web): align environment tables with usage density * feat(web): add search to every environment table * fix(web): preserve unavailable provider recovery actions * fix(web): resolve default provider enabled state * fix(web): guard provider edits and search displayed table values * test(web): exercise provider restore permission with saved account
* feat(web): build the schedules destination and sidebar * fix(web): align schedules layout with Paper measurements * fix(web): refine schedules visual states and recovery actions * fix(web): finish schedules browser visual review * fix(web): preserve schedule failure diagnostics in history
* feat(server): prompt agents to suggest session orchestration * feat(mobile): expose session orchestration settings * fix(mobile): clarify orchestration setting availability
* fix(web): improve desktop update controls and release notes * fix(web): clarify downloading release notes and reuse update helpers
Expose the existing environment setting in the shared desktop and web General panel, with search targeting and reset support. Document where to find it. Validation: 26 targeted settings tests, web typecheck, targeted lint and format checks, and production web build passed. No interactive UI verification performed.
…g a dead reconnect A paired client's session lasts 30 days and there is no renewal path, so every remote environment eventually hits expiry. The server knew the credential had expired, but the wire error flattened that into `invalid_credential`, and the client rendered a generic red "<label> is offline" banner whose Reconnect button re-sent the same dead credential forever. The only visible difference from an ordinary network outage was the banner's colour. Expiry now survives the trip. `EnvironmentAuthInvalidError` carries an optional `credentialExpired` flag, set when the underlying cause is a session or websocket expiry. It is optional and only sent when true, so older clients keep decoding the payload and fall back to `reason`. The clients act on it. `presentConnectionState` exposes `blockedReason`, and the shared `connectionNeedsPairing` helper answers the question both surfaces were getting wrong: retrying cannot clear an authentication block. Web swaps the banner to "<label> needs pairing again" and drops the Reconnect button; mobile does the same and stops promising it "will keep retrying automatically". Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
|
Opened against the wrong repository by mistake — this targets the goodbirdhq/phoenix fork, not upstream. Reopening there. Apologies for the noise. |
|
Macroscope skipped reviewing this pull request. Per-review cost limit exceeded (workspace setting). This review would cost an estimated $63.06, which exceeds your per-review limit of $15.00. The top 3 files driving up this estimate:
Tip To get this pull request reviewed, you can:
|
ApprovabilityVerdict: Not approved Macroscope's review found this PR not approvable — The PR is a broad fork/rebrand and capability expansion touching authentication, agent execution, persistence, desktop/mobile behavior, and production release workflows—not just an isolated credential-message fix. It also changes product defaults and suppresses a static-analysis job, so the combined runtime and operational impact requires human review. Not approved because:
Review your spending limits in Billing settings, or comment |
Problem
A paired client's session lasts 30 days (
DEFAULT_SESSION_TTL) and there is no renewal path, so every remote environment eventually hits expiry. When it does, the server knows exactly what happened —SessionStore.verifyfails withSessionTokenExpiredError— but the wire error flattens that intoinvalid_credential. The client then shows a generic red "<label>is offline" banner whose Reconnect button re-sends the same dead credential forever.The only visible difference from an ordinary network outage is the banner's colour, because
ChatViewrenders the same title for supervisor phaseofflineand phaseblocked. Diagnosing a real instance of this took reading server-side trace logs to discover the word "expired" had been thrown away.Fix
Expiry now survives the trip to the client, and both clients act on it.
EnvironmentAuthInvalidErrorcarries an optionalcredentialExpiredflag. It is an optional key sent only when true, following the existingdpopFailureReasonprecedent, so older clients keep decoding the payload and fall back toreason. Adding a newreasonliteral instead would have broken strict decoding on older clients.serverAuthCredentialExpiredclassifies the underlying cause, covering session and websocket expiry, and everyfailEnvironmentAuthInvalidcall site passes it.presentConnectionStateexposesblockedReason, and a sharedconnectionNeedsPairinghelper answers the question both surfaces were getting wrong: retrying cannot clear an authentication block. A DPoP proof hint is now suppressed for expiry, where it would misdirect.<label>needs pairing again" with a key icon, and the Reconnect button is gone. Connections remains as the way forward.Desktop inherits the web change. The Connections panel improves for free via
connectionStatusText.Testing
vp test runacross the touched scopes: 233 passing. Typecheck clean on contracts, client-runtime, server, web, and mobile; targeted lint and format clean.Tests were written first and watched fail — server classification (expired vs. the revoked near-miss), the client message mapping,
blockedReasonexposure, andconnectionNeedsPairing.Before/after screenshots are not attached, which this repo requires for UI changes. Capturing them means running the app in a browser, which needs maintainer sign-off — happy to produce them on request before merge.
Deliberately out of scope
rotate/renew/refreshanywhere in server auth today; adding one is a new auth-protocol feature, not a bug fix. This PR makes expiry legible and recoverable, it does not stop the 30-day logout. Worth a follow-up, ideally with a warning before expiry while the client can still act.retryNowdoes not re-sample connectivity (supervisor.ts), so Reconnect also silently fails whilenavigator.onLineis latched off after sleep. Real, separate concern, separate PR.Model: Claude Opus 5 (1M context). Harness: Claude Code.
🤖 Generated with Claude Code