Skip to content

Stop Codex collab child watchers leaking after the parent session ends - #551

Merged
alexeyzimarev merged 2 commits into
mainfrom
fix/codex-child-watcher-leak
Aug 13, 2026
Merged

alexeyzimarev merged 2 commits into
mainfrom
fix/codex-child-watcher-leak

Conversation

@alexeyzimarev

Copy link
Copy Markdown
Member

Closes #550 (AI-1928)

Problem

A Codex collab child watcher had no exit path at all — observed as 16 kcap watch … --vendor codex processes from a single session still running two days after the parent session ended, each stuck in an infinite reconnect-and-resend loop against the server:

  1. ShouldEndOnIdle deliberately excluded codex children, assuming "the parent's session-end teardown finalizes them" — but that teardown finalizes server-side records only, never the local processes.
  2. The parent-exit watchdog never arms: children are spawned by the parent session watcher, whose ancestry contains no codex process for GetCodingAgentPid to resolve, so every child starts with "No parent pid supplied; parent-exit watchdog disabled".
  3. KillWatcher was never called from WatchCommand, and the server's StopWatcher signal only reaches the session watcher's connection.

Fix (three layers)

  • Parent teardown stops its children. The session watcher tracks every child watcher key it spawns (codex, and the same-shaped gemini/opencode scans) and stops them on its own way out via the new WatcherManager.KillWatchers — SIGTERM first, so each child runs its final drain before the parent posts session-end; the existing per-child 5s force-kill bound keeps a wedged child from stalling the parent's exit.
  • Reap ceiling as backstop. Codex child watchers join the idle ceiling for the hard-killed-parent orphan case: 6h default on the new KCAP_CODEX_SUBAGENT_REAP_MINUTES knob (KCAP_CODEX_SUBAGENT_IDLE_MINUTES is already the subagent-stop grace), and never while a tool call is in flight — mirroring the Claude subagent ceiling from Reap leaked Claude subagent watchers with an idle ceiling #517. The original "an idle self-exit would end nothing server-side" rationale went stale with [AI-1861] Post live subagent-stop from Codex collab child watchers #526: children post their own live subagent-stop, and once exited the watcher is GONE, which is exactly the state the server sweep can finalize.
  • Re-engagement stays safe. ScanCodexSubagents re-ensures a child watcher whenever its rollout's mtime advances instead of gating on first sight only, so a reaped child whose subagent re-engages is respawned and resumes from the server frontier. Without this, the ceiling would turn re-engagement into silent content loss for the rest of the parent's life.

Testing

  • Unit: codex-child ceiling eligibility/window/knob routing + ChildRolloutAdvanced gate (TDD, watched red first); the codex regression guard that pinned the old ineligible behavior is superseded and updated.
  • Integration: KillWatchers_stops_every_tracked_child_and_clears_their_pid_files — one live and one already-dead child in a batch, both swept.
  • Full unit suite: only the known rotating local timing flakes (daemon/PTY areas, disjoint from this diff). AOT publish: no IL3050/IL2026 warnings.

README documents the new knob and reaping behavior in the same PR.

🤖 Generated with Claude Code

A Codex collab child watcher had no exit path at all: it was deliberately
excluded from the idle ceiling (on the assumption the parent's teardown
finalizes it — that teardown is server-side records only), it gets no
--parent-pid watchdog (its spawner is the parent session watcher, whose
ancestry contains no codex process), and the server's StopWatcher only
reaches the session watcher's connection. Observed: 16 children from one
session still running two days after the parent exited, each reconnecting
and resending the same tail gap forever.

Three layers close the leak:

- The parent session watcher now stops every child watcher it spawned
  (codex/gemini/opencode) as part of its own teardown, via the new
  WatcherManager.KillWatchers — SIGTERM, so each child runs its final
  drain before the parent posts session-end.
- Codex child watchers join the idle ceiling as a leak backstop for a
  hard-killed parent: 6h default on the new KCAP_CODEX_SUBAGENT_REAP_MINUTES
  knob (KCAP_CODEX_SUBAGENT_IDLE_MINUTES is already the stop grace), never
  while a tool call is in flight. The old "an idle self-exit would end
  nothing server-side" rationale went stale when children gained their own
  live subagent-stop POST (#526); once exited, the watcher is GONE and the
  server sweep can finalize it.
- ScanCodexSubagents re-ensures a child watcher whenever its rollout's
  mtime advances instead of gating on first sight only, so a reaped child
  whose subagent re-engages is respawned and resumes from the server
  frontier — without this, the ceiling would turn re-engagement into
  silent content loss for the rest of the parent's life.

Closes #550

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
@qodo-code-review

Copy link
Copy Markdown

PR Summary by Qodo

Stop Codex collab child watcher processes leaking after parent session exit

🐞 Bug fix ✨ Enhancement 🧪 Tests 📝 Documentation 🕐 40+ Minutes

Grey Divider

AI Description

• Track and SIGTERM all spawned child watchers during session watcher teardown.
• Add Codex subagent reap ceiling with new env knob to backstop orphaned children.
• Re-ensure Codex child watchers when rollout mtime advances to prevent silent gaps.
Diagram

graph TD
  A(["WatchCommand (session watcher)"]) --> B(["Scan*Subagents"])
  B --> C(("Child watcher process")) --> D[("PID file")]
  A --> E(["WatcherManager.KillWatchers"])
  E --> C
  A --> F[/"Env config knobs"/]
  F --> A
  subgraph Legend
    direction LR
    _svc(["Service/Module"]) ~~~ _proc(("Process")) ~~~ _cfg[/"Config"/] ~~~ _db[("File store")]
  end
Loading
High-Level Assessment

The following are alternative approaches to this PR:

1. Spawn children in a killable process group/job
  • ➕ OS-enforced cleanup if the parent exits normally
  • ➕ Avoids maintaining an in-memory set of child keys
  • ➖ Platform-specific complexity (Windows job objects vs Unix process groups)
  • ➖ Harder to preserve the 'SIGTERM then bounded force-kill' drain semantics
2. Make StopWatcher server broadcast to child watchers
  • ➕ Centralized lifecycle control; parent teardown becomes less critical
  • ➕ Works even if parent loses local state
  • ➖ Requires protocol/server changes and reliable child registration semantics
  • ➖ Doesn’t cover local-only orphan cases if the connection is already gone
3. Fix parent PID watchdog resolution for children
  • ➕ Makes children self-terminate when parent dies, without relying on teardown
  • ➕ Reduces need for long idle reaping windows
  • ➖ Non-trivial ancestry/PID discovery; brittle across launchers and shells
  • ➖ Still needs a backstop for cases where watchdog can’t arm

Recommendation: Keep the PR’s layered approach (parent teardown kill + Codex child reap ceiling + re-ensure-on-mtime-advance). It directly addresses the observed leak modes without requiring OS- or server-specific changes, while preserving safety (no reap during tool calls) and correctness (respawn on re-engagement to avoid silent transcript loss).

Files changed (5) +258 / -52

Enhancement (1) +11 / -0
WatcherManager.csAdd batch child-watcher kill helper +11/-0

Add batch child-watcher kill helper

• Adds WatcherManager.KillWatchers to concurrently terminate multiple watchers using existing KillWatcher semantics. This enables parent session watchers to reliably stop all spawned children on exit while keeping per-child bounded force-kill behavior.

src/Capacitor.Cli/WatcherManager.cs

Bug fix (1) +116 / -35
WatchCommand.csTrack spawned children, add Codex reap ceiling, and respawn on rollout advance +116/-35

Track spawned children, add Codex reap ceiling, and respawn on rollout advance

• Introduces tracking for all spawned child watcher keys and stops them during session watcher teardown. Makes Codex child watchers eligible for an idle-based self-reap ceiling (new env mapping) with tool-in-flight safety, and updates Codex subagent scanning to re-ensure watchers when the rollout file mtime advances via a new testable gate.

src/Capacitor.Cli/Commands/WatchCommand.cs

Tests (2) +119 / -17
WatcherLifecycleTests.csIntegration test for batch kill sweeping live and stale child pid files +27/-0

Integration test for batch kill sweeping live and stale child pid files

• Adds an integration test ensuring KillWatchers stops a live watcher and also cleans up a pid file pointing to a non-existent process. Verifies the batch operation doesn’t abort on already-dead entries and removes pid files for both.

test/Capacitor.Cli.Tests.Integration/WatcherLifecycleTests.cs

WatchCommandTests.csUnit tests for Codex child reap eligibility, knob parsing, and mtime gate +92/-17

Unit tests for Codex child reap eligibility, knob parsing, and mtime gate

• Replaces the prior regression guard that kept Codex children ineligible for idle exit with tests asserting Codex child watcher eligibility and tool-in-flight safety. Adds coverage for the new KCAP_CODEX_SUBAGENT_REAP_MINUTES mapping/parsing and for ChildRolloutAdvanced behavior (first-sight fire, mtime-only thereafter, per-rollout independence).

test/Capacitor.Cli.Tests.Unit/WatchCommandTests.cs

Documentation (1) +12 / -0
README.mdDocument Codex subagent reap knob and behavior +12/-0

Document Codex subagent reap knob and behavior

• Adds KCAP_CODEX_SUBAGENT_REAP_MINUTES to the configuration table. Documents why Codex collab subagent watchers can leak and explains the new parent-teardown stop plus idle self-reap backstop and respawn-on-reengagement behavior.

README.md

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: 50d41382d4

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

// skips this, which is exactly the orphan case the codex-child reap ceiling backstops.
if (spawnedChildWatcherKeys.Count > 0) {
Log($"Stopping {spawnedChildWatcherKeys.Count} spawned child watcher(s)");
await WatcherManager.KillWatchers(spawnedChildWatcherKeys);

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P1 Badge Avoid killing a reused PID after a child self-reaps

When a Codex child hits the new idle-reap ceiling, RunWatch exits without removing its .pid file. If the OS reuses that PID before the parent exits, this cleanup passes the stale key to KillWatcher, which trusts the PID file and calls Process.Kill without verifying process identity, potentially terminating an unrelated process. The child should remove its own PID state on exit, or cleanup must verify a stored process start identity before killing.

Useful? React with 👍 / 👎.

Copy link
Copy Markdown
Member Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Fixed in 1b8679c. The pid file now carries the incarnation's ProcessStartToken on line 2 (the daemon pid-file layout); KillWatcher spares on a conclusive token mismatch — "ambiguity never kills" — and sweeps the stale file instead. Watchers additionally retire their own pid file on graceful exit (RemoveOwnPidFile, guarded to this incarnation), so the reap-then-teardown window doesn't arise in the first place. Covered by KillWatcher_spares_a_recycled_pid_and_sweeps_the_stale_file (watched fail first: the bystander was killed) and RemoveOwnPidFile_removes_only_this_incarnations_file.

// skips this, which is exactly the orphan case the codex-child reap ceiling backstops.
if (spawnedChildWatcherKeys.Count > 0) {
Log($"Stopping {spawnedChildWatcherKeys.Count} spawned child watcher(s)");
await WatcherManager.KillWatchers(spawnedChildWatcherKeys);

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Badge Terminate child watchers gracefully before awaiting their drain

When the parent exits through StopWatcher, SIGINT, or another path that does not run a vendor teardown with InlineDrainAsync, this call can lose the child's most recent transcript tail. KillWatchers delegates to KillWatcher, whose Process.Kill(entireProcessTree: false) forcibly terminates the process rather than sending SIGTERM, so the child's registered signal handler and final drain/spool code never execute despite the new cleanup relying on them. Send a real graceful signal first, or explicitly inline-drain each child before the force kill.

Useful? React with 👍 / 👎.

Copy link
Copy Markdown
Member Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Fixed in 1b8679c. Confirmed: Process.Kill(entireProcessTree: false) is SIGKILL on Unix — the trap test observed exit 137 pre-fix. KillWatcher now sends SIGTERM first via a libc kill P/Invoke, keeping the existing 5s wait + force-kill fallback, so the child's registered SIGTERM handler runs its final drain + undelivered-tail spool. KillWatcher_sigterms_first_so_the_watcher_can_run_its_drain pins it (bash trap exits 42, not 137/143). Windows keeps the hard stop — no SIGTERM semantics there; recovery stays with the spool/import paths, and the docs now say so.

@qodo-code-review

qodo-code-review Bot commented Aug 13, 2026 •

Copy link
Copy Markdown

Code Review by Qodo

🐞 Bugs (0) 📘 Rule violations (0) 📎 Requirement gaps (0) 🎨 UX issues (0) 🔗 Cross-repo conflicts (0) 📜 Skill insights (0)

Grey Divider


Action required

1. Child watchers hard-killed ✓ Resolved 🐞 Bug ≡ Correctness
Description
WatchCommand now stops spawned child watchers on parent teardown via WatcherManager.KillWatchers,
but KillWatcher immediately uses Process.Kill, so children may not run their SIGTERM/cts.Cancel
shutdown final drain and undelivered-tail spool. This can truncate subagent transcripts right when
the parent exits, contradicting the new teardown’s “final drain before session-end” safety goal.
Code

src/Capacitor.Cli/Commands/WatchCommand.cs[R880-883]

+        if (spawnedChildWatcherKeys.Count > 0) {
+            Log($"Stopping {spawnedChildWatcherKeys.Count} spawned child watcher(s)");
+            await WatcherManager.KillWatchers(spawnedChildWatcherKeys);
+        }
Evidence
The new teardown path explicitly expects SIGTERM-driven graceful shutdown for child watchers, but
the kill implementation uses Process.Kill immediately. The watcher’s graceful shutdown path (signal
handlers canceling cts) is what triggers the shutdown final drain and tail spooling; bypassing it
risks transcript truncation. The repo’s own SIGTERM implementation in ProcessReaper uses
UnixPtyInterop.kill on Unix rather than Process.Kill, indicating Process.Kill is not the intended
SIGTERM mechanism there.

src/Capacitor.Cli/Commands/WatchCommand.cs[875-883]
src/Capacitor.Cli/Commands/WatchCommand.cs[286-317]
src/Capacitor.Cli/Commands/WatchCommand.cs[807-851]
src/Capacitor.Cli/WatcherManager.cs[249-265]
src/Capacitor.Cli.Daemon/Services/ProcessReaper.cs[168-186]

Agent prompt
The issue below was found during a code review. Follow the provided context and guidance below and implement a solution

### Issue description
`WatchCommand` teardown assumes each child watcher gets a graceful stop (so it can run its shutdown final drain + spool), but `WatcherManager.KillWatcher` currently calls `Process.Kill(...)` as the first action. On Unix, the repo’s SIGTERM implementation uses `UnixPtyInterop.kill(..., SIGTERM)` rather than `Process.Kill`, so the current approach can bypass the child watcher’s SIGTERM handlers and skip its drain/spool.

### Issue Context
This PR newly relies on `KillWatchers` as the normal teardown mechanism for spawned child watchers, so the termination semantics now directly impact transcript correctness.

### Fix Focus Areas
- src/Capacitor.Cli/WatcherManager.cs[249-265]
- src/Capacitor.Cli/WatcherManager.cs[287-297]

### What to change
1. Update `KillWatcher` to implement **SIGTERM-first** semantics on Unix (e.g., via a small P/Invoke wrapper like the daemon’s `UnixPtyInterop.kill(pid, SIGTERM)`), then wait up to the existing 5s for graceful exit.
2. Keep the existing bounded **force-kill fallback** (SIGKILL / `Process.Kill(entireProcessTree: true)`), but only after the grace window.
3. Adjust comments/docs in `KillWatcher`/`KillWatchers` and `WatchCommand` teardown to match the actual behavior.
4. Consider mirroring the daemon’s approach for process groups if watchers can spawn descendants that should be reaped with the watcher.

ⓘ Copy this prompt and use it to remediate the issue with your preferred AI generation tools



Remediation recommended

2. Verbose #550 inline comments ✓ Resolved 📘 Rule violation ⚙ Maintainability
Description
The PR adds several long, explanatory inline comment blocks in code and tests that could be
shortened or replaced with clearer naming/extracted helpers. This increases maintenance burden and
makes the code harder to scan over time.
Code

src/Capacitor.Cli/Commands/WatchCommand.cs[R875-878]

+        // #550: stop the child watchers this session watcher spawned — the parent is the only
+        // process that knows they exist, so its teardown is their stop signal on every exit path
+        // (StopWatcher, parent-exit, idle timeout, signals). SIGTERM lets each child run its own
+        // final drain before the parent posts session-end below. A parent killed hard (SIGKILL)
Evidence
PR Compliance ID 5 requires concise comments; the cited regions introduce multi-line narrative
explanations (problem history, rationale, and behavior guarantees) directly in code paths and tests
instead of relying on self-explanatory structure/naming.

CLAUDE.md: Keep comments minimal and prefer self-explanatory code
src/Capacitor.Cli/Commands/WatchCommand.cs[875-883]
src/Capacitor.Cli/Commands/WatchCommand.cs[1391-1405]
src/Capacitor.Cli/WatcherManager.cs[287-296]
test/Capacitor.Cli.Tests.Unit/WatchCommandTests.cs[559-565]

Agent prompt
The issue below was found during a code review. Follow the provided context and guidance below and implement a solution

## Issue description
Several newly-added comments are overly verbose and repeat implementation rationale inline, reducing readability and maintainability.

## Issue Context
PR Compliance ID 5 requires keeping comments minimal and preferring self-explanatory code.

## Fix Focus Areas
- src/Capacitor.Cli/Commands/WatchCommand.cs[670-675]
- src/Capacitor.Cli/Commands/WatchCommand.cs[875-883]
- src/Capacitor.Cli/Commands/WatchCommand.cs[1391-1405]
- src/Capacitor.Cli/WatcherManager.cs[287-296]
- test/Capacitor.Cli.Tests.Unit/WatchCommandTests.cs[559-565]

ⓘ Copy this prompt and use it to remediate the issue with your preferred AI generation tools


Grey Divider

Tip of the day
💡 Did you know, you can type 'qodo, fix this' on a finding and the fix lands right on your PR

More tips ↗ | Customize Qodo ↗ | Qodo docs ↗

Grey Divider

Qodo Logo

Comment thread src/Capacitor.Cli/Commands/WatchCommand.cs Outdated
Comment thread src/Capacitor.Cli/Commands/WatchCommand.cs
Two review findings were real and are now fixed with watched-red tests:

- Process.Kill is SIGKILL on Unix (the old "Send SIGTERM" comment was
  wrong), so a "gracefully stopped" child never ran its final drain +
  undelivered-tail spool. KillWatcher now signals SIGTERM first via a
  libc kill P/Invoke (exit 143->trap proves delivery), keeping the 5s
  force-kill bound. Windows keeps the hard stop; recovery stays with the
  spool/import paths.

- A self-reaped watcher leaves its pid file behind, so the parent's new
  teardown could kill an unrelated process through a recycled pid. The
  pid file now carries the incarnation's ProcessStartToken on line 2
  (daemon pid-file layout); KillWatcher spares on a conclusive token
  mismatch ("ambiguity never kills") and watchers retire their own pid
  file on graceful exit as the first line of defence.

Also trims the more verbose comment blocks flagged by review.

AI-1928 / #550

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
@alexeyzimarev
alexeyzimarev merged commit e5d2f9b into main Aug 13, 2026
6 checks passed
@alexeyzimarev
alexeyzimarev deleted the fix/codex-child-watcher-leak branch August 13, 2026 16:35
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Codex collab child watchers leak indefinitely after the parent session ends

1 participant