IHost.StopAsync() never returns when a Balanced-mode Wolverine node is shut down while holding a large agent load. Found while trying to wire SlowTests into CI for #3779 — this is what blocks it.
Reproduction
docker compose up -d postgresql
dotnet test src/Testing/SlowTests/SlowTests.csproj --framework net9.0 \
--filter 'FullyQualifiedName~agent_reassignment_at_scale'
The test's assertions all pass. DisposeAsync then stops the three hosts in reverse creation order; the first two stop cleanly and the third never returns. dotnet test therefore never reports, and the process has to be killed. Reproduced 4/4, including on a completely idle machine.
The controlled pair
The only change is a bound on the teardown stop in agent_reassignment_at_scale.DisposeAsync:
DisposeAsync |
Result |
await host.StopAsync() (as committed) |
hangs indefinitely; run never reports |
await host.StopAsync().WaitAsync(20.Seconds()) |
passed, 2m18s |
So the hang is entirely in host shutdown. Nothing about the test's own logic is involved — with the stop bounded, the suite is green.
What is known about the wedge
- It is the leader that hangs.
_hosts.Reverse() means the last host stopped is the first one created — node 1, the leader, the node that was running the whole old generation. The other two shut down normally: the assignment table drains 478 → 322 → 192 and then stops moving.
- It is not the simulated slow stop.
DisposeAsync sets _telemetry.StopDelay = TimeSpan.Zero before stopping anything, so the fake agents stop instantly during teardown.
- Nothing is running.
dotnet-stack report against the wedged process shows every managed thread parked — thread-pool workers idle on LowLevelLifoSemaphore.Wait, the main thread blocked in the xUnit entry point, no CPU burn at all (11s of CPU across 40 minutes of wall clock). That rules out a spin or a lock convoy and points at an await that is never completed.
- The runtime goes quiet mid-shutdown. The node's
wolverine_nodes.health_check row stops updating roughly two minutes in, and the row is never deleted — so shutdown gets far enough to stop the health-check loop and never far enough to deregister the node.
Scale that provokes it
agent_reassignment_at_scale runs 3 nodes over a 480-agent universe with the shipped defaults (AgentStartBatchSize = 50, MaxAgentStartParallelism = 10), 240 agents starting on node 1 before two more nodes join. The sibling tests in the same project — agent_assignment_at_scale and slow_starts_outrun_reply_windows — complete normally (3m51s for the pair), so this is not every Balanced-mode shutdown; something about a leader shutting down after a large reassignment wave is the trigger.
Why it matters beyond the test
A host that never completes StopAsync is a deployment that never completes a rolling restart: the orchestrator waits out its termination grace period and then SIGKILLs the pod, which for a durable node means the shutdown path that deregisters the node and releases its agents never runs. That is the same starting condition as the GH-3753 chain.
Blocks
🤖 Generated with Claude Code
https://claude.ai/code/session_0116vfBcKwcjWn8msM4ZjkuA
IHost.StopAsync()never returns when a Balanced-mode Wolverine node is shut down while holding a large agent load. Found while trying to wireSlowTestsinto CI for #3779 — this is what blocks it.Reproduction
The test's assertions all pass.
DisposeAsyncthen stops the three hosts in reverse creation order; the first two stop cleanly and the third never returns.dotnet testtherefore never reports, and the process has to be killed. Reproduced 4/4, including on a completely idle machine.The controlled pair
The only change is a bound on the teardown stop in
agent_reassignment_at_scale.DisposeAsync:DisposeAsyncawait host.StopAsync()(as committed)await host.StopAsync().WaitAsync(20.Seconds())So the hang is entirely in host shutdown. Nothing about the test's own logic is involved — with the stop bounded, the suite is green.
What is known about the wedge
_hosts.Reverse()means the last host stopped is the first one created — node 1, the leader, the node that was running the whole old generation. The other two shut down normally: the assignment table drains478 → 322 → 192and then stops moving.DisposeAsyncsets_telemetry.StopDelay = TimeSpan.Zerobefore stopping anything, so the fake agents stop instantly during teardown.dotnet-stack reportagainst the wedged process shows every managed thread parked — thread-pool workers idle onLowLevelLifoSemaphore.Wait, the main thread blocked in the xUnit entry point, no CPU burn at all (11s of CPU across 40 minutes of wall clock). That rules out a spin or a lock convoy and points at an await that is never completed.wolverine_nodes.health_checkrow stops updating roughly two minutes in, and the row is never deleted — so shutdown gets far enough to stop the health-check loop and never far enough to deregister the node.Scale that provokes it
agent_reassignment_at_scaleruns 3 nodes over a 480-agent universe with the shipped defaults (AgentStartBatchSize = 50,MaxAgentStartParallelism = 10), 240 agents starting on node 1 before two more nodes join. The sibling tests in the same project —agent_assignment_at_scaleandslow_starts_outrun_reply_windows— complete normally (3m51s for the pair), so this is not every Balanced-mode shutdown; something about a leader shutting down after a large reassignment wave is the trigger.Why it matters beyond the test
A host that never completes
StopAsyncis a deployment that never completes a rolling restart: the orchestrator waits out its termination grace period and then SIGKILLs the pod, which for a durable node means the shutdown path that deregisters the node and releases its agents never runs. That is the same starting condition as the GH-3753 chain.Blocks
SlowTestscannot be added to a CI job while one of its tests wedges the runner. Bounding the teardown inside the test would make CI green while hiding this.🤖 Generated with Claude Code
https://claude.ai/code/session_0116vfBcKwcjWn8msM4ZjkuA