Skip to content

fix(lock): verify process identity on macOS before honouring a PID lock (#4210) - #4229

Closed
ntdat812 wants to merge 1 commit into
volcengine:mainfrom
ntdat812:fix/darwin-pid-reuse-lock-4210
Closed

ntdat812 wants to merge 1 commit into
volcengine:mainfrom
ntdat812:fix/darwin-pid-reuse-lock-4210

Conversation

@ntdat812

Copy link
Copy Markdown
Contributor

Fixes #4210.

The gap

The PID-reuse guard from #1088 reads /proc/<pid>/cmdline, and that branch is sys.platform.startswith("linux"). macOS has no /proc, so _is_pid_alive() degraded to a bare os.kill(pid, 0) — which answers "is a process alive", not "is OpenViking alive".

The reported consequence is not a warning, it is a permanent outage: a reboot leaves .openviking.pid behind, the OS hands that PID to something else (Spotlight's mdwrite in the incident), and every start raises DataDirectoryLocked. The impostor never exits, so launchd retried for seven days — ~50,000 crashes, ~330 MB of logs, no path to self-healing.

The change

The identity lookup moves into _process_cmdline(pid), with the Linux /proc read unchanged and a Darwin branch that asks ps:

/bin/ps -p <pid> -o command=

_is_pid_alive() then applies the same rule it already applied on Linux: a live PID whose command line does not mention openviking is a recycled PID, not a lock holder.

The fail-safe direction is deliberate and tested: when the command line cannot be read — ps missing, timing out, exiting non-zero, or answering nothing — _process_cmdline() returns None and the caller keeps trusting the liveness probe. Not being able to identify a process never causes its lock to be stolen.

I dropped the redundant and "openviking-server" not in cmdline from the comparison: openviking-server contains openviking, so the second test could never change the outcome.

Two existing tests changed, and why

test_live_pid_blocks_acquisition and test_error_message_includes_remediation write PID 1 to the lock file and expect it to block. Since #1088 that outcome has depended on the platform, because PID 1 is init/launchd, not OpenViking:

  • Windows: measured failing on main before this PR — os.kill(1, 0) raises, _is_pid_alive returns False, no exception is raised (1 failed, 4 passed).
  • Linux: /proc/1/cmdline is init/systemd, which the Stale PID lock file causes cascading Gateway failure: unhandled rejections + session deadlock #1088 check rejects — so it can only pass where /proc/1/cmdline is unreadable and the except OSError: pass fallback runs.
  • macOS: passes today only because no identity check exists; this PR would make it fail there too.

Both tests now stub _is_pid_alive, so they test the lock semantics they are named for instead of the host's process table. That is a change to existing tests, so I am flagging it rather than burying it — if you would rather I leave them as they are, the alternative is a skipif(sys.platform != "linux") on both, and I can switch.

Tests

tests/misc/test_process_lock_pid_reuse.py — 10 tests. They monkeypatch sys.platform and subprocess.run, so the macOS path is covered from Linux and Windows CI as well; there is no Darwin-only skip.

Coverage: the recycled-PID case (a live non-OpenViking process no longer holds the lock), the real-instance case (still blocks), the exact ps argv and timeout, four "cannot tell" outcomes that must keep the lock, the end-to-end path through acquire_data_dir_lock(), and a guard that no other platform shells out.

Against main with only openviking/utils/process_lock.py reverted: 10 failed. With the change: 15 passed across both lock test files (5 pre-existing + 10 new).

ruff check and ruff format --check clean on the three touched files. (ruff check tests/misc/ also reports an unrelated pre-existing I001 in test_retrieval_enable_intent.py, untouched here.)

…ck (volcengine#4210)

The PID-reuse guard added in volcengine#1088 reads /proc/<pid>/cmdline, which exists
only on Linux. On macOS `_is_pid_alive()` fell back to a bare os.kill(pid, 0),
so after a reboot left a stale .openviking.pid behind and the OS handed that
PID to an unrelated process, every start raised DataDirectoryLocked and the
server never recovered — the reporter measured ~50,000 launchd retries over
seven days against a PID held by Spotlight's mdwrite.

Move the identity lookup into `_process_cmdline()` and add a Darwin branch
that asks `ps`. When the command line cannot be read the function returns
None and the caller keeps trusting the liveness probe, so an unreadable
process never causes a lock to be stolen.

`test_live_pid_blocks_acquisition` and its sibling pinned PID 1, whose
verdict now depends on the identity check on macOS as well as Linux (and
which already failed on Windows, where os.kill(1, 0) does not succeed). Both
now stub the liveness check so they test the lock semantics they are named
for.
@qin-ctx

qin-ctx commented Sep 14, 2026

Copy link
Copy Markdown
Collaborator

Thanks for the patch and the reproduction work. For #4210, we have decided to keep recovery from leftover PID locks after an unclean shutdown as an operator responsibility. When startup cannot establish that a lock is stale, retaining the lock and refusing startup is acceptable. The operator must verify that no OpenViking instance is using the data directory before clearing the leftover lock.

We are closing this PR without merging because we are not expanding automatic recovery with platform-specific command-line identity checks. This is a scope decision; we acknowledge the startup failure described in the issue.

@qin-ctx qin-ctx closed this Sep 14, 2026
@github-project-automation github-project-automation Bot moved this from Backlog to Done in OpenViking project Sep 14, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

Status: Done

Development

Successfully merging this pull request may close these issues.

macOS: stale PID lock misjudged as alive after reboot due to PID reuse — server crash-loops and never self-heals

2 participants