Skip to content

Balanced-mode host shutdown hangs indefinitely when the node holds a large number of agents #3781

Description

@jeremydmiller

IHost.StopAsync() never returns when a Balanced-mode Wolverine node is shut down while holding a large agent load. Found while trying to wire SlowTests into CI for #3779 — this is what blocks it.

Reproduction

docker compose up -d postgresql
dotnet test src/Testing/SlowTests/SlowTests.csproj --framework net9.0 \
  --filter 'FullyQualifiedName~agent_reassignment_at_scale'

The test's assertions all pass. DisposeAsync then stops the three hosts in reverse creation order; the first two stop cleanly and the third never returns. dotnet test therefore never reports, and the process has to be killed. Reproduced 4/4, including on a completely idle machine.

The controlled pair

The only change is a bound on the teardown stop in agent_reassignment_at_scale.DisposeAsync:

DisposeAsync Result
await host.StopAsync() (as committed) hangs indefinitely; run never reports
await host.StopAsync().WaitAsync(20.Seconds()) passed, 2m18s

So the hang is entirely in host shutdown. Nothing about the test's own logic is involved — with the stop bounded, the suite is green.

What is known about the wedge

  • It is the leader that hangs. _hosts.Reverse() means the last host stopped is the first one created — node 1, the leader, the node that was running the whole old generation. The other two shut down normally: the assignment table drains 478 → 322 → 192 and then stops moving.
  • It is not the simulated slow stop. DisposeAsync sets _telemetry.StopDelay = TimeSpan.Zero before stopping anything, so the fake agents stop instantly during teardown.
  • Nothing is running. dotnet-stack report against the wedged process shows every managed thread parked — thread-pool workers idle on LowLevelLifoSemaphore.Wait, the main thread blocked in the xUnit entry point, no CPU burn at all (11s of CPU across 40 minutes of wall clock). That rules out a spin or a lock convoy and points at an await that is never completed.
  • The runtime goes quiet mid-shutdown. The node's wolverine_nodes.health_check row stops updating roughly two minutes in, and the row is never deleted — so shutdown gets far enough to stop the health-check loop and never far enough to deregister the node.

Scale that provokes it

agent_reassignment_at_scale runs 3 nodes over a 480-agent universe with the shipped defaults (AgentStartBatchSize = 50, MaxAgentStartParallelism = 10), 240 agents starting on node 1 before two more nodes join. The sibling tests in the same project — agent_assignment_at_scale and slow_starts_outrun_reply_windows — complete normally (3m51s for the pair), so this is not every Balanced-mode shutdown; something about a leader shutting down after a large reassignment wave is the trigger.

Why it matters beyond the test

A host that never completes StopAsync is a deployment that never completes a rolling restart: the orchestrator waits out its termination grace period and then SIGKILLs the pod, which for a durable node means the shutdown path that deregisters the node and releases its agents never runs. That is the same starting condition as the GH-3753 chain.

Blocks

🤖 Generated with Claude Code

https://claude.ai/code/session_0116vfBcKwcjWn8msM4ZjkuA

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions