Skip to content

aspire run --start-debug-session hangs while stopping an orphaned AppHost #19269

Description

@adamint

Is there an existing issue for this?

Describe the bug

An AppHost launched by the VS Code extension survived after its owning aspire run --start-debug-session CLI exited. Later aspire run and aspire ps commands found its auxiliary socket, connected at the transport level, and then hung indefinitely waiting for the AppHost to answer the initial RPC handshake.

The final visible CLI log entry was:

[dbug] BundleNuGetPackageCache: NuGet search returned 177986 bytes

That search had completed and was not the blocker. The stuck CLI had no child process and never reached AppHost build or launch.

Observed timeline

  • 21:40: aspire run --start-debug-session --nologo started as PID 16380.
  • 21:48: PID 16380 logged a termination signal and Exit code: 0.
  • Its AppHost PID 16554 remained active under PPID 1, along with DCP and application resources.
  • 22:41: a new VS Code extension host launched aspire ps --follow --format json --nologo; it connected to the existing AppHost.
  • Three later aspire run --start-debug-session processes and a non-following aspire ps --format json process all remained stuck.
  • The existing AppHost, DCP, API, worker, and all five later CLI processes were still alive more than an hour after the original launch.

Example process state:

16554     1 S   dotnet .../DotnetProject.AppHost.dll
16603     1 Ss  dcp start-apiserver --monitor 16554 ...
16604 16603 Ss  dcp run-controllers ...
17237 16941 Ss  aspire ps --follow --format json --nologo
17502 16941 S   aspire run --start-debug-session --nologo
19764 16941 Ss  aspire ps --format json --nologo
19912 16941 S   aspire run --start-debug-session --nologo
20505 16941 S   aspire run --start-debug-session --nologo --debug

The later run logs show DotNetAppHostProject connecting to the old auxiliary socket, but no response from the first RPC. The non-following ps log similarly stops after connecting to the same socket.

Relevant code path

RunCommand.ExecuteAsync awaits FindAndStopRunningInstanceAsync before launching. Its comment says a failed stop should not block AppHost startup, but this chain has no bounded handshake:

RunCommand.ExecuteAsync
  -> DotNetAppHostProject.FindAndStopRunningInstanceAsync
  -> RunningInstanceManager.StopRunningInstanceAsync
  -> AppHostAuxiliaryBackchannel.ConnectAsync
  -> CreateFromSocketAsync
  -> GetAppHostInformationAsync / GetCapabilitiesAsync

If the socket accepts but the AppHost never answers JSON-RPC, the command does not reach the intended StopFailed result and startup never continues.

There also appears to be a preceding lifecycle problem: CliOrphanDetector was active and had the original CLI PID plus stable start time, but the AppHost did not terminate after that PID exited.

Expected behavior

  • A foreground AppHost should terminate when its owning CLI exits, or be collected as orphaned.
  • An unhealthy auxiliary socket should not hang aspire run, aspire ps, or aspire stop indefinitely.
  • The initial RPC handshake and stop attempt should be bounded. A failed existing-instance stop should return StopFailed and preserve the flow documented in RunCommand.

Environment

  • Aspire CLI: 13.5.0+25c30ef12379ea5802c74bdd7fd7cf63351d9115
  • CLI build ID: 13.500.26.41106
  • Channel: staging
  • AppHost source SHA: f5f526ec248e49a84875cfe1671594bf27c1d145
  • AppHost version: 13.6.0-dev
  • macOS 26.6 (25G72), Apple Silicon
  • .NET SDK 10.0.201
  • Started through the VS Code extension with --start-debug-session

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    Type

    Projects

    No projects

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions