Skip to content

[AI-78] Refuse foreground kapacitor agent start when another daemon is alive - #51

Merged
alexeyzimarev merged 3 commits into
mainfrom
alexeyzimarev/ai-78-foreground-pid-guard
May 8, 2026
Merged

alexeyzimarev merged 3 commits into
mainfrom
alexeyzimarev/ai-78-foreground-pid-guard

Conversation

@alexeyzimarev

Copy link
Copy Markdown
Member

Summary

  • CLI half of the AI-78 fix. The detached path (-d/--detach) has guarded against duplicate launches via the PID file since day one (StartDetached line 67 — Agent daemon already running (PID …). Use \kapacitor agent stop` first.). Foreground was the only hole. A second kapacitor agent start— run by mistake, by an automation, or by a parallel terminal — happily spawns a freshkapacitor-daemon. Cold-start fresh daemons call DaemonConnectwith emptylive_agents(the orchestrator hasn't been constructed yet,GetLiveAgentIdsis null, the?? []fallback fires), the server silently replaces the active daemon's(owner, name)slot viaDaemonRegistry.Register, then ReconcileDaemon([])` flips every running hosted agent to Failed. The displaced real daemon keeps its WebSocket open and never notices.
  • The fix: mirror the detached PID-file check into StartForegroundAsync, write the PID file when the foreground daemon launches so subsequent starts in any mode see it, and clean it up on exit. IsOurDaemon's StartTime check already handles the recycled-PID case if the parent dies hard.
  • Help text + README updated to document the new behaviour, per the project convention of keeping help-usage.txt / help-<cmd>.txt / README in sync.

Tested locally:

$ kapacitor agent --help        # new "Notes:" section visible
$ kapacitor agent start         # exits 1, message: "Agent daemon already running (PID 54115)…"

Independent of the server-side belt (kurrent-io/Kurrent.Capacitor PR 590) which tightens ReconcileDaemon to ignore empty live_agents. Either alone closes the cascade; both together make it double-safe.

See AI-78 for the full investigation, including macOS process-log evidence of 5 distinct kapacitor-daemon PIDs today and confirmation the real daemon's HubConnection never reconnected.

Test plan

  • AOT publish clean (no IL3050/IL2026 warnings).
  • Smoke test: kapacitor agent start with the existing daemon alive exits 1 with the expected message.
  • Help text renders the new "Notes:" section correctly.
  • Reviewer: confirm the PID file is properly cleaned up on Ctrl+C in foreground mode (the try/finally deletes it; verified manually but worth a second look).

🤖 Generated with Claude Code

… is alive

The detached path (`-d`/`--detach`) has guarded against duplicate launches
via the PID file since day one. The foreground path didn't, so a second
`kapacitor agent start` (run by mistake, by an automation, by a parallel
terminal) would happily spawn a fresh kapacitor-daemon process. That
fresh process calls DaemonConnect with empty live_agents (orchestrator
not yet wired → GetLiveAgentIds returns []), the server's Register
silently replaces the active daemon's slot, ReconcileDaemon([]) flips
every running hosted agent to Failed, and the displaced real daemon
keeps its WebSocket open without ever noticing.

Mirror the detached path's PID-file check into StartForegroundAsync,
write the PID file when the foreground daemon launches (so subsequent
starts in any mode see it), and clean it up on exit. IsOurDaemon's
StartTime check already handles the recycled-PID case if the parent
dies hard and leaves a stale file.

Server-side belt: kurrent-io/Kurrent.Capacitor PR 590 also tightens
the ReconcileDaemon guard to ignore empty live_agents, so the cascade
becomes impossible regardless of what produced the second
DaemonConnect. Either fix alone closes the bug; both together make it
double-safe.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
@linear

linear Bot commented May 8, 2026

Copy link
Copy Markdown

AI-78

@qodo-code-review

Copy link
Copy Markdown

Review Summary by Qodo

Prevent foreground daemon start when another daemon is alive

🐞 Bug fix ✨ Enhancement

Grey Divider

Walkthroughs

Description
• Add PID file guard to foreground daemon start mode
• Prevent duplicate daemon launches that silently displace active daemons
• Mirror detached mode's existing PID check into foreground path
• Update help text and README to document new behavior
Diagram
flowchart LR
  A["kapacitor agent start"] --> B{"Check PID file<br/>in foreground mode"}
  B -->|Daemon alive| C["Exit with error<br/>PID already running"]
  B -->|No daemon| D["Start new daemon<br/>Write PID file"]
  D --> E["Wait for exit"]
  E --> F["Clean up PID file<br/>in finally block"]
  C --> G["Return exit code 1"]
  F --> H["Return exit code"]
Loading

Grey Divider

File Changes

1. src/kapacitor/Commands/AgentCommands.cs 🐞 Bug fix +26/-2

Add PID guard and cleanup to foreground daemon start

• Add PID file existence check at start of StartForegroundAsync to refuse launch if daemon already
 running
• Write PID file when foreground daemon starts so subsequent starts in any mode detect it
• Wrap WaitForExitAsync in try/finally block to ensure PID file cleanup on exit
• Add detailed comments explaining the AI-78 bug and the fix rationale

src/kapacitor/Commands/AgentCommands.cs


2. README.md 📝 Documentation +3/-1

Document foreground daemon duplicate prevention behavior

• Update agent stop description to clarify it stops daemon in any mode
• Add new paragraph explaining that agent start refuses to launch when daemon already exists
• Document the historical issue of concurrent daemons silently failing hosted agents

README.md


3. src/Kapacitor.Core/Resources/help-agent.txt 📝 Documentation +7/-1

Add notes section documenting daemon start guard

• Update stop subcommand description to indicate it works for foreground or background daemons
• Add new "Notes:" section explaining the duplicate daemon prevention behavior
• Document the requirement to use kapacitor agent stop before starting a new daemon

src/Kapacitor.Core/Resources/help-agent.txt


View more (1)
4. src/Kapacitor.Core/Resources/help-usage.txt 📝 Documentation +2/-2

Update help text for daemon start and stop commands

• Update agent start description to note it refuses launch if daemon already running
• Update agent stop description to clarify it works for both foreground and background modes

src/Kapacitor.Core/Resources/help-usage.txt


Grey Divider

Qodo Logo

@qodo-code-review

qodo-code-review Bot commented May 8, 2026 •

Copy link
Copy Markdown

Code Review by Qodo

🐞 Bugs (1) 📘 Rule violations (0) 📎 Requirement gaps (0)

Grey Divider


Action required

1. PID file deleted incorrectly ✓ Resolved 🐞 Bug ☼ Reliability
Description
StartForegroundAsync always deletes agent.pid on exit, even if another concurrent agent start
overwrote the PID file to point at a different still-running daemon. This can orphan the running
daemon (stop/status no longer find it) and re-open the duplicate-daemon hole the PID guard is meant
to close.
Code

src/kapacitor/Commands/AgentCommands.cs[R74-87]

+        // Write the PID file so the guard above (and `kapacitor agent stop` /
+        // `agent status`) sees the foreground daemon too. Best-effort cleanup
+        // on exit; if the parent dies hard, IsOurDaemon's StartTime check
+        // handles the recycled-PID case so a stale file doesn't lock anyone
+        // out.
+        WritePidFile(process);
+
+        try {
+            await process.WaitForExitAsync();

-        return process.ExitCode;
+            return process.ExitCode;
+        } finally {
+            try { File.Delete(PidPath); } catch { /* best-effort */ }
+        }
Evidence
The foreground path writes agent.pid and later unconditionally deletes it; because PID writes
overwrite existing content, a concurrent start can replace the PID file before the first foreground
instance exits, and the first instance then deletes the other daemon’s PID file. The daemon
lifecycle commands rely on the PID file for stop/status behavior, so deleting the wrong PID file
breaks management of the still-running daemon.

src/kapacitor/Commands/AgentCommands.cs[34-88]
src/kapacitor/Commands/AgentCommands.cs[90-134]
src/kapacitor/Commands/AgentCommands.cs[182-203]
src/kapacitor/Commands/AgentCommands.cs[136-178]

Agent prompt
The issue below was found during a code review. Follow the provided context and guidance below and implement a solution

### Issue description
`StartForegroundAsync` deletes `agent.pid` unconditionally in `finally`. If another concurrent start overwrote the PID file after we wrote it, the `finally` block deletes the *other* daemon’s PID file, orphaning that daemon and defeating the PID guard.

### Issue Context
- Foreground path now calls `WritePidFile(process)` and then `File.Delete(PidPath)` in `finally`.
- `WritePidFile` uses `File.WriteAllText`, which overwrites existing content.

### Fix Focus Areas
- Make PID file cleanup conditional: before deleting, re-read `agent.pid` and only delete if it still matches the process we started (PID and, if present, StartTicks).
- Consider making PID acquisition more atomic to reduce TOCTOU (e.g., create-new semantics / lock file / re-check right before writing and abort/kill the newly spawned daemon if another live daemon won the race).

### Fix Focus Areas (code pointers)
- src/kapacitor/Commands/AgentCommands.cs[34-88]
- src/kapacitor/Commands/AgentCommands.cs[182-203]
- src/kapacitor/Commands/AgentCommands.cs[90-134]

ⓘ Copy this prompt and use it to remediate the issue with your preferred AI generation tools



Remediation recommended

2. Status parses PID file wrong 🐞 Bug ≡ Correctness
Description
kapacitor status reads the entire agent.pid file and parses it as an int; when WritePidFile
records StartTicks on a second line, int.TryParse fails and status reports an invalid PID file
even while the agent is running. This becomes more likely now that foreground start also writes the
PID file.
Code

src/kapacitor/Commands/AgentCommands.cs[R74-80]

+        // Write the PID file so the guard above (and `kapacitor agent stop` /
+        // `agent status`) sees the foreground daemon too. Best-effort cleanup
+        // on exit; if the parent dies hard, IsOurDaemon's StartTime check
+        // handles the recycled-PID case so a stale file doesn't lock anyone
+        // out.
+        WritePidFile(process);
+
Evidence
WritePidFile can write a two-line PID file (pid\nstartTicks). StatusCommand reads the full
file and tries to parse it as a single integer, which fails for multi-line content. With this PR,
foreground mode now writes the PID file too, so users running the daemon in foreground are much more
likely to see kapacitor status report an invalid PID file.

src/kapacitor/Commands/AgentCommands.cs[182-203]
src/kapacitor/Commands/StatusCommand.cs[43-63]
src/kapacitor/Commands/AgentCommands.cs[34-88]

Agent prompt
The issue below was found during a code review. Follow the provided context and guidance below and implement a solution

### Issue description
`StatusCommand` reads the entire `agent.pid` content and passes it to `int.TryParse`. When the PID file contains both PID and StartTicks (two lines), parsing fails and `kapacitor status` reports an invalid PID file even though the daemon is running.

### Issue Context
- `AgentCommands.WritePidFile` may write either one line (PID) or two lines (PID + UTC ticks).
- This PR adds PID-file writing to foreground mode, so this status bug becomes visible in foreground usage.

### Fix Focus Areas
- Update `StatusCommand` to parse only the first non-empty line as PID (ignore any subsequent lines).
- (Optional) Reuse a shared PID-file parsing helper to avoid drift between `agent status` and `status`.

### Fix Focus Areas (code pointers)
- src/kapacitor/Commands/StatusCommand.cs[43-63]
- src/kapacitor/Commands/AgentCommands.cs[182-221]

ⓘ Copy this prompt and use it to remediate the issue with your preferred AI generation tools


Grey Divider

Qodo Logo

Comment thread src/kapacitor/Commands/AgentCommands.cs
#1: StartForegroundAsync's finally block deleted agent.pid unconditionally,
which orphaned a concurrent legitimate daemon's PID file in this race:

  1. Process A's daemon-A exits cleanly. Process A enters finally.
  2. Process B's `kapacitor agent start` reads the PID file, sees PID-A;
     IsOurDaemon returns false (PID-A's process is gone), guard passes.
  3. Process B spawns daemon-B, writes PID-B.
  4. Process A's finally deletes the PID file — orphaning daemon-B.

Now we re-read agent.pid in finally and only delete if it still matches
the process we spawned. Belt-and-braces against PID race losses.

#2: StatusCommand had its own PID-file parser that did
ReadAllText().Trim() + int.TryParse(...) — which fails on the multi-line
PID|StartTicks format AgentCommands.WritePidFile produces. Without the
foreground guard this only triggered for `-d` daemons; this PR makes
foreground daemons hit it too. Parse the first non-empty line instead.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
@alexeyzimarev

Copy link
Copy Markdown
Member Author

Thanks Qodo. Both fixed in 532000d:

#1 PID file deleted incorrectly — Real race. Fixed by re-reading agent.pid in the finally and only deleting if it still matches the process we spawned. The window is: our daemon exits cleanly → a concurrent kapacitor agent start passes the guard (because IsOurDaemon correctly returns false on our exited PID) → it writes its own PID-B → our finally deletes the file → daemon-B is orphaned. Now if (ReadPidFile() is { } current && current.Pid == process.Id) File.Delete(...).

I'm leaving the more atomic acquisition (lock file / create-new) as a follow-up; it's a bigger refactor and the conditional-delete closes the window without changing the file-format contract. The remaining TOCTOU between guard-read and WritePidFile is benign on the new behavior — the worst case is the second daemon overwrites our PID file (we then no-op the delete, and the second daemon's stop/status keeps working).

#2 Status parses PID file wrong — Fixed. StatusCommand now takes the first non-empty line of agent.pid instead of trying to parse the whole multi-line PID\nStartTicks content as one int. Bug existed before this PR too (any -d daemon would also hit it) but the foreground guard makes it user-visible for foreground users, which is what Qodo flagged.

I considered consolidating into a single PID-file reader (Qodo's "shared helper" suggestion) and decided against it for this PR — AgentCommands.ReadPidFile returns a private PidEntry record with StartTicks, which StatusCommand doesn't need; promoting it costs more API surface than the bug warrants. Worth doing if/when we add a third reader.

Copilot AI left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull request overview

This PR closes the foreground-mode gap in the CLI that allowed launching a second kapacitor-daemon while another daemon was already running, which could disrupt hosted agents on the server. It also updates help/README documentation to reflect the new behavior and fixes PID-file parsing in kapacitor status.

Changes:

  • Add a foreground agent start guard against an already-running daemon and write/cleanup the PID file for foreground runs.
  • Fix StatusCommand PID parsing to handle the newer multi-line PID file format.
  • Update CLI help text and README to document the refusal behavior and that agent stop applies to both modes.

Reviewed changes

Copilot reviewed 5 out of 5 changed files in this pull request and generated 1 comment.

Show a summary per file
File Description
src/kapacitor/Commands/StatusCommand.cs Parse only the first PID-file line (supports PID + optional StartTicks format).
src/kapacitor/Commands/AgentCommands.cs Refuse duplicate foreground starts; write PID file for foreground; best-effort cleanup on exit.
src/Kapacitor.Core/Resources/help-usage.txt Update top-level usage text for new start/stop behavior.
src/Kapacitor.Core/Resources/help-agent.txt Update agent subcommand help and add “Notes” about refusal behavior.
README.md Document that agent stop applies to both modes and that agent start refuses duplicates.
Comments suppressed due to low confidence (1)

src/kapacitor/Commands/StatusCommand.cs:64

  • StatusCommand now supports the two-line PID file format, but it still only checks whether any process exists with that PID. Since the PID file optionally contains StartTicks to protect against PID reuse, this can incorrectly report "running" when the PID has been recycled to an unrelated process. Consider parsing the optional second line and validating Process.StartTime (or reusing the same IsOurDaemon-style logic as AgentCommands) so status output stays correct under PID recycling.
            // The PID file is one or two lines: PID, optionally followed by
            // process StartTicks (UTC ticks of Process.StartTime). Parse only
            // the first non-empty line as the PID — naively passing the whole
            // contents to int.TryParse fails on the two-line format and
            // mis-reports a running daemon as "invalid PID file".
            var firstLine = (await File.ReadAllTextAsync(pidPath))
                .Split('\n', StringSplitOptions.RemoveEmptyEntries | StringSplitOptions.TrimEntries)
                .FirstOrDefault();

            if (int.TryParse(firstLine, out var pid)) {
                try {
                    System.Diagnostics.Process.GetProcessById(pid);
                    await Console.Out.WriteLineAsync($"running (PID {pid})");
                } catch (ArgumentException) {
                    await Console.Out.WriteLineAsync("not running (stale PID file)");
                }

💡 Add Copilot custom instructions for smarter, more guided reviews. Learn how to get started.

Comment thread src/kapacitor/Commands/AgentCommands.cs Outdated
Comment on lines +35 to +46
// Refuse to start when an existing kapacitor-daemon is alive — even in
// foreground mode. Without this guard, a second `kapacitor agent start`
// (run by mistake, by an automation, by a parallel terminal) would
// happily connect to the server and call DaemonConnect with an empty
// live_agents list, which used to mass-fail every hosted agent owned
// by the original daemon (AI-78). The detached path has had this
// guard since day one; foreground was the only hole.
if (ReadPidFile() is { } existing && IsOurDaemon(existing.Pid, existing.StartTicks)) {
await Console.Error.WriteLineAsync($"Agent daemon already running (PID {existing.Pid}). Use `kapacitor agent stop` first.");

return 1;
}
Previous PID-file guard had a TOCTOU window: two concurrent
`kapacitor agent start` invocations could both observe no live daemon,
both spawn daemons, and only one's PID would end up in the file. The
loser's daemon would then race-write DaemonConnect with empty
live_agents and trip the cascade we're meant to prevent.

Add agent.start.lock opened with FileShare.None (POSIX flock(LOCK_EX),
Windows native sharing constraint). Foreground holds the lock for the
daemon's entire lifetime; detached holds it just for the
check + spawn + WritePidFile window before the parent exits.

Two concurrent starts now serialize at the OS level: the second's
TryAcquireStartLock returns null (IOException from FileShare.None) and
it refuses with "Another `kapacitor agent start` is already in
progress …" The lock is per-open-handle so even SIGKILL on the parent
releases it cleanly.

Refactored StartForegroundAsync into two parts: the lock-acquiring
guard wrapper and the original spawn body (now SpawnForegroundAsync)
to keep the lock-acquire/release scope easy to read.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
@alexeyzimarev

Copy link
Copy Markdown
Member Author

You're right — TOCTOU was real, not benign. Pushed 23a9871 with an actual exclusive lock.

agent.start.lock is opened with FileShare.None — POSIX flock(LOCK_EX) on macOS/Linux, native sharing-mode constraint on Windows. Foreground holds it for the whole daemon lifetime; detached holds it just for the check + spawn + WritePidFile window before the parent exits.

Two concurrent kapacitor agent start invocations now serialize at the OS level: the loser's TryAcquireStartLock returns null and it refuses with Another \kapacitor agent start` is already in progress or holds the daemon lock.` The lock is per-open-handle, so SIGKILL on the parent releases it cleanly — no zombie locks.

Refactored StartForegroundAsync into a thin lock-acquiring wrapper plus a new SpawnForegroundAsync that holds the original spawn body, to keep the acquire/release scope obvious.

Re: CI — the test failure (ImportChainsAsync_dispatches_independent_chains_in_parallel) is a known-flaky timing assertion (spreadMs < 160 after a 200ms-delayed parallel HTTP fan-out). It hit main itself on May 6 (run 25422621258) with the same assertion. Rerun passed; would be worth bumping the threshold or relaxing the test in a separate PR.

@alexeyzimarev
alexeyzimarev merged commit bfdecdb into main May 8, 2026
4 checks passed
@alexeyzimarev
alexeyzimarev deleted the alexeyzimarev/ai-78-foreground-pid-guard branch May 8, 2026 14:22
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants