Summary
On Windows, a second logon session of the same user starts a second ADE brain supervisor. The two supervisors then fight over the brain's named pipe. In the observed run, the user's working brain was stopped and restarted, and one supervisor stayed in an endless socket_owned_by_other restart loop.
Any second session of the same Windows user can trigger this: a Remote Desktop child session, fast user switching, or a normal Remote Desktop logon.
Where it happened
- Windows 11 Pro 26200, packaged ADE (stable channel), user
arul2.
- 2026-09-29, during a spike for per-lane Windows computer use. The spike created a Remote Desktop child session (session 2) next to the console session (session 1).
Timeline (UTC, from %USERPROFILE%\.ade\runtime\brain-service-29103ebb0490.ps1.log)
- Session 2 logged on. The
HKCU\...\Run entry started a supervisor in session 2 (pid 23032), which started a brain (pid 39584). Its serve failed or competed with the session 1 brain. (I stopped both session 2 processes by hand at about 15:43 to protect the real brain. This probably left brain-service-*.ps1.pid.json pointing at a dead supervisor.)
- 15:45:57: an ADE desktop app opened in session 2 (a
.txt file association chose ADE). It ran the brain-service install. The launcher script brain-service-29103ebb0490.ps1 was rewritten at that time.
- 15:46:00: the original brain (pid 45868, uptime 22.6 h) exited with code 1.
supervisor=28532 brain exited pid=45868 exitCode=1 lifetimeMs=81308850.
- 15:46:01: the install started a new supervisor (57216). The old supervisor (28532) was still alive and restarted its brain at the same moment (pid 5832).
- The old supervisor's brain won the pipe. The new supervisor's brains failed every time with backoff 1 s, 2 s, 4 s, 8 s ...:
ADE brain socket is already in use ... \\.\pipe\ade-runtime-stable-23b8c33b935af5fb ... Another ADE brain is accepting connections on this named pipe. (last-failure.json, code socket_owned_by_other).
- Recovery by hand: I stopped the old supervisor 28532 and brain 5832. The new supervisor started brain 51564, which is healthy (fresh
heartbeat.json).
Root cause (from the code)
apps/ade-cli/src/serviceManager/installWindows.ts registers the supervisor under WINDOWS_RUN_KEY (HKCU\Software\Microsoft\Windows\CurrentVersion\Run). A Run entry runs at every logon of the user, so each session gets its own supervisor. Nothing makes the supervisor single-instance per user.
- The supervisor PID record (
<launcher>.pid.json, written in windowsSupervisor.ts) holds one supervisor and is last-writer-wins. A second supervisor overwrites it.
queryWindowsSupervisor() and the install / repair path (removeWindowsRunEntryIfPresent, installWindowsServiceImpl) find the predecessor only through that record. A live supervisor that is not in the record is invisible, so install starts a new supervisor next to it and never stops it.
- When a supervisor's brain fails with
socket_owned_by_other, the supervisor keeps retrying with backoff instead of detecting that another supervisor of the same channel owns the brain.
Expected
One supervisor and one brain per user and channel, whatever the number of logon sessions. A second session must not start, stop, or replace the brain of the first session.
Suggested fix
- Make the supervisor single-instance per user and channel: hold a named mutex that is visible across sessions, for example
Global\ade-supervisor-<channel>-<user SID>. A supervisor that cannot take it exits at once and logs why. (Local\ names are per session, so they do not work here.)
- Make predecessor detection in install / repair independent of the PID record: find supervisors by command line (the launcher path) in all sessions of the user, and stop all of them before starting the new one.
- When the brain fails with
socket_owned_by_other, the supervisor should check whether another supervisor of the channel owns the pipe and then exit, instead of looping.
- Decide whether the brain may run from a non-console session at all. A child session or RDP session ends when the user signs out of it, and the brain should not live there.
Test ideas
- Unit: two supervisors started for the same channel → the second exits and logs the mutex owner.
- Unit: install with a stale PID record and a live untracked supervisor → the untracked supervisor is stopped.
- Manual on Windows 11 Pro: enable child sessions (
WTSEnableChildSessions), open a child session, confirm only one ADE.exe ... serve exists across both sessions, then open ADE desktop inside the child session and confirm the session 1 brain keeps its pid.
Related
Found during the per-lane Windows computer-use spike (Windows Desktop, the Windows counterpart of Mac Desktop).
Summary
On Windows, a second logon session of the same user starts a second ADE brain supervisor. The two supervisors then fight over the brain's named pipe. In the observed run, the user's working brain was stopped and restarted, and one supervisor stayed in an endless
socket_owned_by_otherrestart loop.Any second session of the same Windows user can trigger this: a Remote Desktop child session, fast user switching, or a normal Remote Desktop logon.
Where it happened
arul2.Timeline (UTC, from
%USERPROFILE%\.ade\runtime\brain-service-29103ebb0490.ps1.log)HKCU\...\Runentry started a supervisor in session 2 (pid 23032), which started a brain (pid 39584). Itsservefailed or competed with the session 1 brain. (I stopped both session 2 processes by hand at about 15:43 to protect the real brain. This probably leftbrain-service-*.ps1.pid.jsonpointing at a dead supervisor.).txtfile association chose ADE). It ran the brain-service install. The launcher scriptbrain-service-29103ebb0490.ps1was rewritten at that time.supervisor=28532 brain exited pid=45868 exitCode=1 lifetimeMs=81308850.ADE brain socket is already in use ... \\.\pipe\ade-runtime-stable-23b8c33b935af5fb ... Another ADE brain is accepting connections on this named pipe.(last-failure.json, codesocket_owned_by_other).heartbeat.json).Root cause (from the code)
apps/ade-cli/src/serviceManager/installWindows.tsregisters the supervisor underWINDOWS_RUN_KEY(HKCU\Software\Microsoft\Windows\CurrentVersion\Run). A Run entry runs at every logon of the user, so each session gets its own supervisor. Nothing makes the supervisor single-instance per user.<launcher>.pid.json, written inwindowsSupervisor.ts) holds one supervisor and is last-writer-wins. A second supervisor overwrites it.queryWindowsSupervisor()and the install / repair path (removeWindowsRunEntryIfPresent,installWindowsServiceImpl) find the predecessor only through that record. A live supervisor that is not in the record is invisible, so install starts a new supervisor next to it and never stops it.socket_owned_by_other, the supervisor keeps retrying with backoff instead of detecting that another supervisor of the same channel owns the brain.Expected
One supervisor and one brain per user and channel, whatever the number of logon sessions. A second session must not start, stop, or replace the brain of the first session.
Suggested fix
Global\ade-supervisor-<channel>-<user SID>. A supervisor that cannot take it exits at once and logs why. (Local\names are per session, so they do not work here.)socket_owned_by_other, the supervisor should check whether another supervisor of the channel owns the pipe and then exit, instead of looping.Test ideas
WTSEnableChildSessions), open a child session, confirm only oneADE.exe ... serveexists across both sessions, then open ADE desktop inside the child session and confirm the session 1 brain keeps its pid.Related
Found during the per-lane Windows computer-use spike (Windows Desktop, the Windows counterpart of Mac Desktop).