fix(coding-agent): self-heal stale worker registrations on resume - #852
Conversation
|
Fixed in 610e038. Reclaim now verifies process identity via |
|
Addressed both findings in 019058c:
|
|
Both addressed in 8ed4aa6 by collapsing reclaim onto the existing finalizer instead of adding coordination:
|
There was a problem hiding this comment.
Cursor Bugbot has reviewed your changes and found 1 potential issue.
❌ Bugbot Autofix is OFF. To automatically fix reported issues with cloud agents, enable autofix in the Cursor dashboard.
Reviewed by Cursor Bugbot for commit 8ed4aa6. Configure here.
|
Fixed in c641a1e. When the process is confirmed dead but the finalizer outlasts the bounded wait, reclaim now throws a retryable error ("still being cleaned up; retry shortly") instead of returning false — so the resume path can never fall through to |
0ac3dfc to
9de578b
Compare
Pull Request is not mergeable
Pull Request is not mergeable
Pull Request is not mergeable
9de578b to
52a604d
Compare
Reopening a saved session used to fail forever when a stopped worker left a tombstoned registration behind (stop timed out, process died later, and finalization was interrupted). The supervisor now detects such stale registrations during create/resume, completes the interrupted stop, and launches a fresh worker for the same saved transcript.
…gistrations A recycled pid used to make a dead worker look alive, so its stale registration was never reclaimed and resume kept failing. Reclaim now checks processStartId, treating a recycled pid as gone; the stop path never signals a pid whose identity no longer matches.
…tion A timed-out stop already has a background finalizer completing the same cleanup, so the resume-time reclaim now awaits it instead of running a duplicate stop that could repeat archival and cron-lock cleanup. Also document the intentional finalizer/reclaim race in the end-to-end test: both paths are covered deterministically by unit tests.
…ervable Align resume-time reclaim with the directional identity verdicts: only a confirmed-gone or confirmed-replaced pid is reclaimed; a transient identity lookup failure leaves the registration untouched.
…inalizer Reclaim now checks confirmed process death first and then delegates the cleanup to scheduleWorkerStopFinalization, so concurrent resumes share one stop instead of duplicating archival, and the bounded wait keeps a resume request from blocking on a finalizer that cannot settle.
…esume When the bounded reclaim wait expires before the finalizer finishes, the resume now fails with a retry hint instead of falling through to reuseWorkerForCreate with a registration whose process is confirmed dead - the exact failure mode this PR heals.
52a604d to
6e3414b
Compare
A descriptor persisted before identity tracking has no processStartId. When supervisor adoption finds such a worker tombstoned and still alive, stopWorker resolved its identity as unknown and refused to send SIGTERM or SIGKILL, and the background finalizer likewise never escalates without a captured start id -- the live process ran forever while its registration could never be finalized or resumed. Before the adoption stop, authenticate on the worker socket (the same proof of ownership the adopt and recovery paths already rely on) and persist the start id observed while the process was alive; the connected client also gives stopWorker its graceful IPC shutdown path. If the handshake fails the identity stays untrusted and the stop keeps its conservative wait-only behavior.
|
Valid finding — fixed in 0517f98. For a pre-identity descriptor (no The fix captures identity the same way the adopt and recovery paths already do: authenticate on the worker's socket (proof the pid is our worker, not a recycled process), then persist the Tests: identity captured+persisted before the stop runs (order asserted), and a failed connect leaves |
Patch release. Bug fixes, small UX additions behind existing surfaces, and a dependency consolidation; no breaking changes, per the no-major-releases policy. Contents since v0.7.1: - #838 in-place queue editing (Alt+Up/Alt+Down browse, Enter/Alt+Enter apply) and queue preservation on interrupt - #850/#851/#852 worker lifecycle truthfulness, timed-out stop finalization, and stale-registration self-heal (the "Session worker is not connected" family) - #1226 Down Arrow stays in a nonempty prompt until the cursor reaches the end - #767 independent expand/collapse for tool calls, a2a messages, and thinking - #1135 agents view keeps expansion state when leaving and returning - #647 login URL copy action - #521 privacy-safe agent analytics with disclosure and opt-out - #846 Homebrew ownership preserved on self-update - #772 sent a2a messages show only message text when expanded - #632 consolidated dependency updates (undici 7.29, biome 2.5.5, marked 18, typescript 7 dev-only, typebox 1.3, aws-sdk bedrock, vitest 4.1.10, et al.) - #1132 stale Gemini test model update (test-only) Missing changelog entries for #838/#850/#851/#852 are added under 0.7.2. Lockstep bump across the root package and the four published packages; example and private workspaces untouched. Lockfile updated surgically (version fields and inter-package ranges only).
…imeIntellect-ai#852) * fix(coding-agent): self-heal stale worker registrations on resume Reopening a saved session used to fail forever when a stopped worker left a tombstoned registration behind (stop timed out, process died later, and finalization was interrupted). The supervisor now detects such stale registrations during create/resume, completes the interrupted stop, and launches a fresh worker for the same saved transcript. * fix(coding-agent): verify process identity before reclaiming stale registrations A recycled pid used to make a dead worker look alive, so its stale registration was never reclaimed and resume kept failing. Reclaim now checks processStartId, treating a recycled pid as gone; the stop path never signals a pid whose identity no longer matches. * fix(coding-agent): defer resume reclaim to an in-flight stop finalization A timed-out stop already has a background finalizer completing the same cleanup, so the resume-time reclaim now awaits it instead of running a duplicate stop that could repeat archival and cron-lock cleanup. Also document the intentional finalizer/reclaim race in the end-to-end test: both paths are covered deterministically by unit tests. * fix(coding-agent): leave reclaim alone when process identity is unobservable Align resume-time reclaim with the directional identity verdicts: only a confirmed-gone or confirmed-replaced pid is reclaimed; a transient identity lookup failure leaves the registration untouched. * fix(coding-agent): route resume reclaim through the single-flighted finalizer Reclaim now checks confirmed process death first and then delegates the cleanup to scheduleWorkerStopFinalization, so concurrent resumes share one stop instead of duplicating archival, and the bounded wait keeps a resume request from blocking on a finalizer that cannot settle. * fix(coding-agent): never hand a confirmed-dead registration back to resume When the bounded reclaim wait expires before the finalizer finishes, the resume now fails with a retry hint instead of falling through to reuseWorkerForCreate with a registration whose process is confirmed dead - the exact failure mode this PR heals. * refactor(coding-agent): use the pid-based identity helper in resume reclaim * fix(coding-agent): capture live worker identity before an adoption stop A descriptor persisted before identity tracking has no processStartId. When supervisor adoption finds such a worker tombstoned and still alive, stopWorker resolved its identity as unknown and refused to send SIGTERM or SIGKILL, and the background finalizer likewise never escalates without a captured start id -- the live process ran forever while its registration could never be finalized or resumed. Before the adoption stop, authenticate on the worker socket (the same proof of ownership the adopt and recovery paths already rely on) and persist the start id observed while the process was alive; the connected client also gives stopWorker its graceful IPC shutdown path. If the handshake fails the identity stays untrusted and the stop keeps its conservative wait-only behavior.
Patch release. Bug fixes, small UX additions behind existing surfaces, and a dependency consolidation; no breaking changes, per the no-major-releases policy. Contents since v0.7.1: - PrimeIntellect-ai#838 in-place queue editing (Alt+Up/Alt+Down browse, Enter/Alt+Enter apply) and queue preservation on interrupt - PrimeIntellect-ai#850/PrimeIntellect-ai#851/PrimeIntellect-ai#852 worker lifecycle truthfulness, timed-out stop finalization, and stale-registration self-heal (the "Session worker is not connected" family) - PrimeIntellect-ai#1226 Down Arrow stays in a nonempty prompt until the cursor reaches the end - PrimeIntellect-ai#767 independent expand/collapse for tool calls, a2a messages, and thinking - PrimeIntellect-ai#1135 agents view keeps expansion state when leaving and returning - PrimeIntellect-ai#647 login URL copy action - PrimeIntellect-ai#521 privacy-safe agent analytics with disclosure and opt-out - PrimeIntellect-ai#846 Homebrew ownership preserved on self-update - PrimeIntellect-ai#772 sent a2a messages show only message text when expanded - PrimeIntellect-ai#632 consolidated dependency updates (undici 7.29, biome 2.5.5, marked 18, typescript 7 dev-only, typebox 1.3, aws-sdk bedrock, vitest 4.1.10, et al.) - PrimeIntellect-ai#1132 stale Gemini test model update (test-only) Missing changelog entries for PrimeIntellect-ai#838/PrimeIntellect-ai#850/PrimeIntellect-ai#851/PrimeIntellect-ai#852 are added under 0.7.2. Lockstep bump across the root package and the four published packages; example and private workspaces untouched. Lockfile updated surgically (version fields and inter-package ranges only).

Problem
If a dead worker's registration did get left behind (stop timed out, the process died later, and the supervisor restarted before cleaning up), you could never reopen that session. Resume kept matching the dead registration and failing with "Session worker is not connected".
Fix
When you open or resume a session and the matching worker turns out to be marked for stop, disconnected, and dead, the supervisor now finishes the old cleanup, drops the stale registration, and starts a fresh worker for the same saved transcript.
What this does not change
Healthy workers are never touched. Neither are workers that are still shutting down or in the middle of recovery.
Validation
npm run checkpassed.Part 3 of 3, on top of #851.
Note
Medium Risk
Changes daemon supervisor worker lifecycle and session resume paths; mistakes could strand sessions or signal wrong PIDs, but identity checks and bounded waits limit blast radius.
Overview
Fixes a case where a stopped worker’s registration could outlive the process (stop timeout, late death, supervisor restart) and block reopening the saved transcript on resume.
On create/resume, when a session path matches a worker that is tombstoned and disconnected, the supervisor now calls
reclaimStaleWorkerRegistration: it only proceeds if process identity is gone or replaced, then drives the existing stop finalizer (boundedSTALE_RECLAIM_WAIT_MSwait) instead of treating the slot as still active. Healthy, connected, or still-stopping workers are unchanged.During supervisor adoption of a stop-pending worker with no
processStartId, it connects to the live worker socket, captures and persists the start id when verification succeeds, sostopWorkerand background finalization can signal the right process instead of waiting on a possibly recycled pid.Validation adds unit coverage for reclaim, adoption identity capture, and concurrent reclaims, plus an E2E path: failed stop, process dies, resume succeeds with a new worker.
Reviewed by Cursor Bugbot for commit 0517f98. Bugbot is set up for automated code reviews on this repo. Configure here.
Note
Fix stale worker registrations blocking session resume in coding agent daemon
stopWorkeris now identity-aware: it captures the entry PID/start-ID and avoids signaling recycled or unknown PIDs. On timeout, it leaves a tombstone and schedules background finalization viafinalizeTimedOutWorkerStop.reclaimStaleWorkerRegistrationprobes process identity (current/replaced/gone/unknown) and either reclaims the slot or raises a retryable error if cleanup is still in progress.processIdExists,isZombieProcess,isProcessAlive) in child-process.ts replace ad-hoc liveness checks and exclude zombie processes from alive counts.Changes since #852 opened
Macroscope summarized 52a604d.