Skip to content

fix(coding-agent): self-heal stale worker registrations on resume - #852

Merged
snimu merged 8 commits into
mainfrom
snimu/worker-resume-self-heal
Aug 11, 2026
Merged

snimu merged 8 commits into
mainfrom
snimu/worker-resume-self-heal

Conversation

@snimu

@snimu snimu commented Aug 7, 2026 •

Copy link
Copy Markdown
Contributor

Problem

If a dead worker's registration did get left behind (stop timed out, the process died later, and the supervisor restarted before cleaning up), you could never reopen that session. Resume kept matching the dead registration and failing with "Session worker is not connected".

Fix

When you open or resume a session and the matching worker turns out to be marked for stop, disconnected, and dead, the supervisor now finishes the old cleanup, drops the stale registration, and starts a fresh worker for the same saved transcript.

What this does not change

Healthy workers are never touched. Neither are workers that are still shutting down or in the middle of recovery.

Validation

  • 3 unit tests: stale registration gets cleaned up; a still-running or healthy worker is left alone.
  • An end-to-end test that replays the original incident: stop fails, process dies, and resume just works — same session, new worker, attach succeeds.
  • npm run check passed.

Part 3 of 3, on top of #851.


Note

Medium Risk
Changes daemon supervisor worker lifecycle and session resume paths; mistakes could strand sessions or signal wrong PIDs, but identity checks and bounded waits limit blast radius.

Overview
Fixes a case where a stopped worker’s registration could outlive the process (stop timeout, late death, supervisor restart) and block reopening the saved transcript on resume.

On create/resume, when a session path matches a worker that is tombstoned and disconnected, the supervisor now calls reclaimStaleWorkerRegistration: it only proceeds if process identity is gone or replaced, then drives the existing stop finalizer (bounded STALE_RECLAIM_WAIT_MS wait) instead of treating the slot as still active. Healthy, connected, or still-stopping workers are unchanged.

During supervisor adoption of a stop-pending worker with no processStartId, it connects to the live worker socket, captures and persists the start id when verification succeeds, so stopWorker and background finalization can signal the right process instead of waiting on a possibly recycled pid.

Validation adds unit coverage for reclaim, adoption identity capture, and concurrent reclaims, plus an E2E path: failed stop, process dies, resume succeeds with a new worker.

Reviewed by Cursor Bugbot for commit 0517f98. Bugbot is set up for automated code reviews on this repo. Configure here.

Note

Fix stale worker registrations blocking session resume in coding agent daemon

  • When resuming a saved session, the daemon supervisor now checks if a previously stopped worker's process is dead or recycled. If so, it reclaims the stale registration immediately rather than blocking the resume.
  • stopWorker is now identity-aware: it captures the entry PID/start-ID and avoids signaling recycled or unknown PIDs. On timeout, it leaves a tombstone and schedules background finalization via finalizeTimedOutWorkerStop.
  • reclaimStaleWorkerRegistration probes process identity (current/replaced/gone/unknown) and either reclaims the slot or raises a retryable error if cleanup is still in progress.
  • Supervisor shutdown no longer blocks on workers that exceed stop deadlines; those workers are tombstoned and finalized asynchronously.
  • New process-probing utilities (processIdExists, isZombieProcess, isProcessAlive) in child-process.ts replace ad-hoc liveness checks and exclude zombie processes from alive counts.

Changes since #852 opened

  • Added self-healing logic for legacy worker descriptors during supervisor adoption [0517f98]
  • Added test coverage for legacy worker descriptor adoption scenarios [0517f98]

Macroscope summarized 52a604d.

Comment thread packages/coding-agent/src/modes/daemon/daemon-supervisor.ts Outdated
Comment thread packages/coding-agent/src/modes/daemon/daemon-supervisor.ts
@snimu

snimu commented Aug 7, 2026

Copy link
Copy Markdown
Contributor Author

Fixed in 610e038. Reclaim now verifies process identity via processStartId (same pattern as recoverWorker): a recycled pid counts as gone, so the stale registration is still reclaimed, and the identity-aware stop path never signals the pid's new owner. Added a regression test with a live pid whose processStartId no longer matches.

Comment thread packages/coding-agent/test/daemon-supervisor-process.test.ts
@snimu

snimu commented Aug 7, 2026

Copy link
Copy Markdown
Contributor Author

Addressed both findings in 019058c:

  • Duplicate stop: reclaim now awaits an in-flight stopFinalization instead of starting a second stopWorker, so archive and cron-lock cleanup can never run twice. New unit test covers the deferral.
  • E2E race: kept intentionally — whichever cleanup wins (background finalizer or resume-time reclaim), the user-visible guarantee under test is that resume succeeds; both individual paths are covered deterministically by unit tests. Documented this in the test.

Comment thread packages/coding-agent/src/modes/daemon/daemon-supervisor.ts Outdated
Comment thread packages/coding-agent/src/modes/daemon/daemon-supervisor.ts Outdated
@snimu

snimu commented Aug 7, 2026

Copy link
Copy Markdown
Contributor Author

Both addressed in 8ed4aa6 by collapsing reclaim onto the existing finalizer instead of adding coordination:

  • Never-settling wait: reclaim now fails fast — it checks for a confirmed-dead process first and returns immediately for live/unknown identities; the wait on the finalizer is bounded (10s), so a resume request always returns promptly.
  • Duplicate cleanup: reclaim no longer calls stopWorker itself. It delegates to scheduleWorkerStopFinalization, which is single-flighted per worker, so concurrent resumes share one stop and archival/cron-lock cleanup cannot run twice. New regression test races two concurrent reclaims and asserts a single stop.

@cursor cursor Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Cursor Bugbot has reviewed your changes and found 1 potential issue.

Fix All in Cursor

❌ Bugbot Autofix is OFF. To automatically fix reported issues with cloud agents, enable autofix in the Cursor dashboard.

Reviewed by Cursor Bugbot for commit 8ed4aa6. Configure here.

Comment thread packages/coding-agent/src/modes/daemon/daemon-supervisor.ts Outdated
@snimu

snimu commented Aug 7, 2026

Copy link
Copy Markdown
Contributor Author

Fixed in c641a1e. When the process is confirmed dead but the finalizer outlasts the bounded wait, reclaim now throws a retryable error ("still being cleaned up; retry shortly") instead of returning false — so the resume path can never fall through to reuseWorkerForCreate with a dead registration. Regression test covers the slow-cleanup timeout.

alexzhang13
alexzhang13 previously approved these changes Aug 11, 2026
stack merge was automatically disabled August 11, 2026 05:11

Pull Request is not mergeable

stack merge was automatically disabled August 11, 2026 07:58

Pull Request is not mergeable

stack merge was automatically disabled August 11, 2026 07:59

Pull Request is not mergeable

@snimu
snimu force-pushed the snimu/worker-resume-self-heal branch from 9de578b to 52a604d Compare August 11, 2026 08:02
Base automatically changed from snimu/worker-shutdown-finalize to main August 11, 2026 08:13
snimu added 5 commits August 11, 2026 10:13
Reopening a saved session used to fail forever when a stopped worker left
a tombstoned registration behind (stop timed out, process died later, and
finalization was interrupted). The supervisor now detects such stale
registrations during create/resume, completes the interrupted stop, and
launches a fresh worker for the same saved transcript.
…gistrations

A recycled pid used to make a dead worker look alive, so its stale
registration was never reclaimed and resume kept failing. Reclaim now
checks processStartId, treating a recycled pid as gone; the stop path
never signals a pid whose identity no longer matches.
…tion

A timed-out stop already has a background finalizer completing the same
cleanup, so the resume-time reclaim now awaits it instead of running a
duplicate stop that could repeat archival and cron-lock cleanup. Also
document the intentional finalizer/reclaim race in the end-to-end test:
both paths are covered deterministically by unit tests.
…ervable

Align resume-time reclaim with the directional identity verdicts: only a
confirmed-gone or confirmed-replaced pid is reclaimed; a transient
identity lookup failure leaves the registration untouched.
…inalizer

Reclaim now checks confirmed process death first and then delegates the
cleanup to scheduleWorkerStopFinalization, so concurrent resumes share
one stop instead of duplicating archival, and the bounded wait keeps a
resume request from blocking on a finalizer that cannot settle.
snimu added 2 commits August 11, 2026 10:13
…esume

When the bounded reclaim wait expires before the finalizer finishes, the
resume now fails with a retry hint instead of falling through to
reuseWorkerForCreate with a registration whose process is confirmed
dead - the exact failure mode this PR heals.
@snimu
snimu force-pushed the snimu/worker-resume-self-heal branch from 52a604d to 6e3414b Compare August 11, 2026 08:13
Comment thread packages/coding-agent/src/modes/daemon/daemon-supervisor.ts
A descriptor persisted before identity tracking has no processStartId. When
supervisor adoption finds such a worker tombstoned and still alive,
stopWorker resolved its identity as unknown and refused to send SIGTERM or
SIGKILL, and the background finalizer likewise never escalates without a
captured start id -- the live process ran forever while its registration
could never be finalized or resumed.

Before the adoption stop, authenticate on the worker socket (the same proof
of ownership the adopt and recovery paths already rely on) and persist the
start id observed while the process was alive; the connected client also
gives stopWorker its graceful IPC shutdown path. If the handshake fails the
identity stays untrusted and the stop keeps its conservative wait-only
behavior.
@snimu

snimu commented Aug 11, 2026

Copy link
Copy Markdown
Contributor Author

Valid finding — fixed in 0517f98.

For a pre-identity descriptor (no processStartId) with a live pid and an intentional-stop tombstone, stopWorker resolved identity as unknown and (correctly) refused to SIGTERM/SIGKILL by pid alone, and the finalizer's stoppedCanSignal stayed false for the same reason — so the live worker ran forever and the registration could never be finalized.

The fix captures identity the same way the adopt and recovery paths already do: authenticate on the worker's socket (proof the pid is our worker, not a recycled process), then persist the processStartId observed while it was alive, before calling stopWorker. The connected client also gives the stop its graceful IPC shutdown path. If the handshake fails, the identity stays untrusted and the stop keeps its conservative wait-only behavior — we never signal a pid whose identity the handshake did not confirm.

Tests: identity captured+persisted before the stop runs (order asserted), and a failed connect leaves processStartId unset with no persist.

@snimu
snimu merged commit 14d6e74 into main Aug 11, 2026
17 checks passed
@snimu
snimu deleted the snimu/worker-resume-self-heal branch August 11, 2026 08:35
@snimu snimu mentioned this pull request Aug 11, 2026
9 tasks done
sethkarten pushed a commit that referenced this pull request Aug 11, 2026
Patch release. Bug fixes, small UX additions behind existing surfaces, and a
dependency consolidation; no breaking changes, per the no-major-releases
policy.

Contents since v0.7.1:
- #838 in-place queue editing (Alt+Up/Alt+Down browse, Enter/Alt+Enter apply)
  and queue preservation on interrupt
- #850/#851/#852 worker lifecycle truthfulness, timed-out stop finalization,
  and stale-registration self-heal (the "Session worker is not connected"
  family)
- #1226 Down Arrow stays in a nonempty prompt until the cursor reaches the end
- #767 independent expand/collapse for tool calls, a2a messages, and thinking
- #1135 agents view keeps expansion state when leaving and returning
- #647 login URL copy action
- #521 privacy-safe agent analytics with disclosure and opt-out
- #846 Homebrew ownership preserved on self-update
- #772 sent a2a messages show only message text when expanded
- #632 consolidated dependency updates (undici 7.29, biome 2.5.5, marked 18,
  typescript 7 dev-only, typebox 1.3, aws-sdk bedrock, vitest 4.1.10, et al.)
- #1132 stale Gemini test model update (test-only)

Missing changelog entries for #838/#850/#851/#852 are added under 0.7.2.

Lockstep bump across the root package and the four published packages;
example and private workspaces untouched. Lockfile updated surgically
(version fields and inter-package ranges only).
0oAstro pushed a commit to 0oAstro/fulcrum that referenced this pull request Aug 11, 2026
…imeIntellect-ai#852)

* fix(coding-agent): self-heal stale worker registrations on resume

Reopening a saved session used to fail forever when a stopped worker left
a tombstoned registration behind (stop timed out, process died later, and
finalization was interrupted). The supervisor now detects such stale
registrations during create/resume, completes the interrupted stop, and
launches a fresh worker for the same saved transcript.

* fix(coding-agent): verify process identity before reclaiming stale registrations

A recycled pid used to make a dead worker look alive, so its stale
registration was never reclaimed and resume kept failing. Reclaim now
checks processStartId, treating a recycled pid as gone; the stop path
never signals a pid whose identity no longer matches.

* fix(coding-agent): defer resume reclaim to an in-flight stop finalization

A timed-out stop already has a background finalizer completing the same
cleanup, so the resume-time reclaim now awaits it instead of running a
duplicate stop that could repeat archival and cron-lock cleanup. Also
document the intentional finalizer/reclaim race in the end-to-end test:
both paths are covered deterministically by unit tests.

* fix(coding-agent): leave reclaim alone when process identity is unobservable

Align resume-time reclaim with the directional identity verdicts: only a
confirmed-gone or confirmed-replaced pid is reclaimed; a transient
identity lookup failure leaves the registration untouched.

* fix(coding-agent): route resume reclaim through the single-flighted finalizer

Reclaim now checks confirmed process death first and then delegates the
cleanup to scheduleWorkerStopFinalization, so concurrent resumes share
one stop instead of duplicating archival, and the bounded wait keeps a
resume request from blocking on a finalizer that cannot settle.

* fix(coding-agent): never hand a confirmed-dead registration back to resume

When the bounded reclaim wait expires before the finalizer finishes, the
resume now fails with a retry hint instead of falling through to
reuseWorkerForCreate with a registration whose process is confirmed
dead - the exact failure mode this PR heals.

* refactor(coding-agent): use the pid-based identity helper in resume reclaim

* fix(coding-agent): capture live worker identity before an adoption stop

A descriptor persisted before identity tracking has no processStartId. When
supervisor adoption finds such a worker tombstoned and still alive,
stopWorker resolved its identity as unknown and refused to send SIGTERM or
SIGKILL, and the background finalizer likewise never escalates without a
captured start id -- the live process ran forever while its registration
could never be finalized or resumed.

Before the adoption stop, authenticate on the worker socket (the same proof
of ownership the adopt and recovery paths already rely on) and persist the
start id observed while the process was alive; the connected client also
gives stopWorker its graceful IPC shutdown path. If the handshake fails the
identity stays untrusted and the stop keeps its conservative wait-only
behavior.
0oAstro pushed a commit to 0oAstro/fulcrum that referenced this pull request Aug 11, 2026
Patch release. Bug fixes, small UX additions behind existing surfaces, and a
dependency consolidation; no breaking changes, per the no-major-releases
policy.

Contents since v0.7.1:
- PrimeIntellect-ai#838 in-place queue editing (Alt+Up/Alt+Down browse, Enter/Alt+Enter apply)
  and queue preservation on interrupt
- PrimeIntellect-ai#850/PrimeIntellect-ai#851/PrimeIntellect-ai#852 worker lifecycle truthfulness, timed-out stop finalization,
  and stale-registration self-heal (the "Session worker is not connected"
  family)
- PrimeIntellect-ai#1226 Down Arrow stays in a nonempty prompt until the cursor reaches the end
- PrimeIntellect-ai#767 independent expand/collapse for tool calls, a2a messages, and thinking
- PrimeIntellect-ai#1135 agents view keeps expansion state when leaving and returning
- PrimeIntellect-ai#647 login URL copy action
- PrimeIntellect-ai#521 privacy-safe agent analytics with disclosure and opt-out
- PrimeIntellect-ai#846 Homebrew ownership preserved on self-update
- PrimeIntellect-ai#772 sent a2a messages show only message text when expanded
- PrimeIntellect-ai#632 consolidated dependency updates (undici 7.29, biome 2.5.5, marked 18,
  typescript 7 dev-only, typebox 1.3, aws-sdk bedrock, vitest 4.1.10, et al.)
- PrimeIntellect-ai#1132 stale Gemini test model update (test-only)

Missing changelog entries for PrimeIntellect-ai#838/PrimeIntellect-ai#850/PrimeIntellect-ai#851/PrimeIntellect-ai#852 are added under 0.7.2.

Lockstep bump across the root package and the four published packages;
example and private workspaces untouched. Lockfile updated surgically
(version fields and inter-package ranges only).
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants