Repository navigation
Conversation
ApprovabilityVerdict: Not approved Macroscope's review found this PR not approvable — The PR substantially changes production provider-session ownership, concurrency, cleanup, credential revocation, authentication flows, and server shutdown behavior across multiple components. Its size and sensitive auth/credential lifecycle impact exceed the scope of changes suitable for automatic approval. Not approved because:
Review your spending limits in Billing settings, or comment |
|
Note on the three red checks — all reproduce on the base branch (
Checks covering this change are green: Test Server 1–3, Rust, CodeRabbit. |
a38ea46 to
1310472
Compare
1310472 to
b3e7733
Compare
Bugbot is paused — on-demand spend limit reachedBugbot uses usage-based billing for this team and has hit its on-demand spend limit. A team admin can raise the spend limit in the Cursor dashboard, or wait for the next billing cycle to continue. |
1 similar comment
Bugbot is paused — on-demand spend limit reachedBugbot uses usage-based billing for this team and has hit its on-demand spend limit. A team admin can raise the spend limit in the Cursor dashboard, or wait for the next billing cycle to continue. |
a8cc38b to
834edb9
Compare
|
Macroscope skipped reviewing this pull request. Per-PR cost limit exceeded (workspace setting). Reviews on this PR have cost $44.13 so far. This review would add an estimated $9.20, bringing the total to $53.33 — above your per-PR limit of $50.00. Tip To get this pull request reviewed, you can:
|
600a8a5 to
d8c75ec
Compare
2b49d24 to
3b5f8cb
Compare
|
Worker 01 — provider-session cleanup ownership (rebased Eighth base rewrite absorbed (
|
463d0c6 to
e10a1e2
Compare
64bc4d7 to
6f41417
Compare
a5ebe9a to
a62d7ad
Compare
6f41417 to
e312ad0
Compare
juliusmarminge
left a comment
There was a problem hiding this comment.
Reopened so the smaller version can land on this PR. See my comment above for the requested shape: a releasing map of deferreds that open/ensureThread/resumeThread await and repeated close/detach join, keep the existing 30s bound, drop the opening-record/revocation-chain/in-band-drain machinery and the serverRuntimeStartup.ts worker changes (separate PR if still needed), and aim for under ~300 lines including two focused tests. Please force-push the rewrite rather than stacking on the current head.
juliusmarminge
left a comment
There was a problem hiding this comment.
Reopened so the smaller version can land on this PR. See my comment above for the requested shape: a releasing map of deferreds that open/ensureThread/resumeThread await and repeated close/detach join, keep the existing 30s bound, drop the opening-record/revocation-chain/in-band-drain machinery and the serverRuntimeStartup.ts worker changes (separate PR if still needed), and aim for under ~300 lines including two focused tests. Please force-push the rewrite rather than stacking on the current head.
fe4f6ad to
87c67bd
Compare
668a26f to
ff93972
Compare
Preserve the complete PR-owned snapshot at 6b22b7f5cee379310267c716b830db0ce484909b while removing historical upstream merge ancestry. Original PR: pingdotgg#11492. The nine later follow-up commits are replayed separately.
…up failures A terminal detach could win the reacquire race between attachThreadOrReject releasing the thread-attach lock and the resource-creating adapter call reacquiring it, stripping the attachment and its credential claim before the op was admitted. The lock now covers bookkeeping and the adapter call in one hold, and markBusy moves before the attach to avoid a lifecycle-lock cycle. Native close failures were also swallowed by Effect.ignore in the Cursor, Claude, ACP, and OpenCode session finalizers, converting failed cleanup into a clean release over still-owned provider resources. They now propagate (or die inside Effect.addFinalizer slots) so the manager keeps ownership, and a failed startup close retains the MCP credential reservation until cleanup status is known. stopSessions now delegates to the closeInstance barrier so sign-in/out joins in-flight startups and failed cleanups instead of trusting projected status. Generated with [Devin](https://devin.ai) Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>
… marks Astra xhigh follow-ups on the provider-cleanup contract: - Claude interrupt timeouts only drop queryContext when the native close succeeded; a failed close stays tracked so the session scope retries it instead of reporting a clean release over a live CLI. - startTurn and compactThread unwind their busy mark with onExit so interruption while queued on the attach lock cannot pin the session busy forever (catch only covered typed failures). - The attach/detach race regression now sweeps 60 deterministic scheduler offsets through the credential-issuance gate. - Close-propagation tests inject typed Effect.fail errors instead of defects, since Effect.ignore preserves defects either way. Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>
…ttach-race sweep before the gate The onExit unwind for startTurn/compactThread decremented busyCount on any non-success exit, so an interrupt landing before markBusy's increment stole a different holder's mark and idled a live turn. markBusy now takes an acquisition flag set inside the same uninterruptible region as the increment, and the unwind only calls markIdle when this call actually took the mark. The attach-lock sweep aimed the stepping scheduler after opening the credential gate, but Deferred.succeed resumes waiters inline on the synchronous scheduler, so the target never engaged. The sweep now arms before the gate opens and reaches offset 520 — past the ~505-op bookkeeping-to-admission boundary where the unfused release/reacquire gap actually sat. Generated with [Devin](https://devin.ai) Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>
… typecheck The two-branch modify callback inferred the union's first member as the result type, leaving outcome unknown under the repo-wide CI typecheck. Generated with [Devin](https://devin.ai) Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>
…ay outbox before its worker restarts A stalled adapter acquisition runs inside an uninterruptible mask, so an unbounded Fiber.interrupt on the effect worker could pin shutdown before provider-session cleanup ever ran. The stop wait is now bounded by the same 30s bound used for cleanup, at both the shutdown finalizer and the relay-startup failure path. The replay harness restart path also forked the effect worker daemon with its layer before reconcileAfterProcessLoss ran, letting the daemon claim persisted rows that reconciliation then reset to pending. The restart path now builds the layer without the daemon, reconciles, then forks the worker — matching production ordering. Generated with [Devin](https://devin.ai) Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>
…edge teardown forkScoped registers its own unbounded Fiber.interrupt finalizer, and it lands after the shutdown teardown finalizer, so it runs first on scope close — a stalled worker still pinned teardown before the bounded wait could run. The worker is now forked detached and bound by a bounded-interrupt finalizer that clears workerFiberRef, keeping a single 30s bound before providerSessions.shutdown. The replay restart path also now honours a caller's runEffectWorker: false instead of forking the daemon unconditionally after reconciliation. Generated with [Devin](https://devin.ai) Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>
A detached worker was forked before workerFiberRef recorded it and its bounded-interrupt finalizer was registered, so an interrupt landing between the fork and the record orphaned a worker parked at activation — free to run against a closed runtime once activation released. The fork, ownership record, and finalizer registration now run uninterruptibly, so every spawned worker is either fully owned or never created. The regression sweeps hold offsets around the handoff and asserts the worker can never run after scope close plus activation release; it fails on the non-atomic version. Generated with [Devin](https://devin.ai) Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>
decorateRuntime carried injectHistory through the spread unguarded, so a runtime a release or replacement already retired could still invoke native thread/inject_items — e.g. a Codex history injection racing a close during ContextHandoffDelivery's pending-write suspension could land on a dying session yet still be durably marked injected. Decorate it like snapshot/rollback: live-runtime checks around the activity touch, then admission into inflightAdapterOps so a release drains the call before scope close and the ownership recheck refuses stale or replaced runtimes, plus the detached-thread refusal when the provider thread carries an appThreadId. Regression coverage proves a released or same-id-replaced runtime rejects injectHistory before the adapter runs (invocation counter stays 0) while a live runtime is admitted. Model: SWE-2 Max (Devin/Cognition) via Cursor harness in T3 Code.
ff93972 to
25c1917
Compare
Summary
Ports #9806 onto
t3code/codex-turn-mapping.A V2 provider session removed its session-map entry before scope cleanup finished, so a replacement
opencould start a new process while the old one still owned credentials and resources — and a hung cleanup silently unblocked reuse.opens and runtime thread attachments (ensureThread/resumeThread/forkThread/startTurn) for its recorded threads are blocked.close/detach/closeInstancecalls join the same cleanupDeferredrather than racing it; adetachthat loses its session to a concurrent close joins the pending cleanup instead of reporting success over still-running work.errorprojection state without claiming the process stopped; successful cleanup clears the block, failed cleanup keeps blocking until restart.Post-review hardening vs. #9806: the runtime attachment gate runs outside
observeActivity(which swallows bookkeeping failures) and outsidestartTurn's busy/idle catch (a rejected attach must not decrement the peer's busy count).Test plan
9fcc9a4bdb:vp test run apps/server/src/orchestration-v2/ProviderSessionManager.test.ts apps/server/src/provider/Layers/ProviderAuthService.test.ts --maxWorkers=1— 169 tests passed. Auth tests preserve upstream shared-binding replacement/exclusion cases and verify a failed peer close prevents invalidation and logout.vp test run apps/server/src/provider/Layers/ProviderAuthService.test.ts apps/server/src/orchestration-v2/ProviderSessionManager.test.ts apps/server/src/orchestration-v2/Adapters/ClaudeAdapterV2.test.ts apps/server/src/orchestration-v2/Adapters/CursorAdapterV2.test.ts apps/server/src/orchestration-v2/Adapters/OpenCodeAdapterV2.test.ts apps/server/src/orchestration-v2/Adapters/AcpAdapterV2.test.ts apps/server/src/mcp/OrchestratorMcpToolkit.integration.test.ts apps/server/src/orchestration-v2/testkit/OrchestratorReplayRecovery.integration.test.ts apps/server/src/serverRuntimeStartup.test.ts --maxWorkers=2— 472 tests passed and the original scheduler sweep timed out; all seven adjacent suites passed (307 tests). The revised core-suite result is above.pnpm run typecheckinapps/serverpassed; targeted formatting check passed.ProviderSessionReleaseErrorinstead of overwriting its reason. The stalled-attach regression reproduced before the fix; the three focused detach timeout cases pass afterward.SWE-2 (Devin) via T3 Code / Cursor harness.
Coordination trace: T3 thread
063442b9-45e8-4f0e-bc47-edea523ce7d0; campaign saphid/t3code-personal#298; original implementation #9806.