Skip to content

[Hermes][Sandbox] managed gateway restart can health-timeout after the supervised gateway exits #7484

Description

@ericksoa

Problem

With NemoClaw 0.0.92, a managed Hermes gateway restart can stop the existing gateway cleanly but never produce a healthy replacement:

nemohermes fbroggini-hermes gateway restart
Failure layer: health timeout - gateway restart failed for fbroggini-hermes. GATEWAY_HEALTH_TIMEOUT

The gateway log shows the old process receiving SIGTERM, draining immediately, stopping the API server, and exiting with code 1 so a supervisor can revive it. Its parent is the expected managed supervisor:

parent_name=bash parent_cmdline="bash /usr/local/bin/nemoclaw-start"
Gateway stopped by an unexpected signal — persisting gateway_state=running
Exiting with code 1 (signal-initiated shutdown without restart request) so systemd Restart=on-failure can revive the gateway.

No replacement becomes healthy before the host-side restart controller times out.

Scope and current evidence

Confirmed:

There is also a latent supervisor defect worth fixing and regression-testing: refresh_hermes_supervised_child_pids() does not explicitly return success. Under set -e, an absent optional dashboard log-tail process can make mark_hermes_gateway_stopped() terminate the supervisor before replacement launch. This is a plausible failure path, but the available log does not prove it caused this particular incident; preparation, launch, gateway health, auxiliary readiness, or crash quarantine could also fail while producing the same outer timeout.

Acceptance criteria

  • Optional gateway/dashboard log bookkeeping explicitly succeeds when an optional tail is absent and restores missing log tails when appropriate.
  • The same managed supervisor remains alive while the old gateway PID is replaced by a new healthy PID.
  • Restart failure evidence distinguishes runtime preparation, launch failure, supervisor loss, gateway-health failure, auxiliary-health failure, and crash quarantine.
  • Failure/debug output includes supervisor and gateway PIDs plus a bounded, redacted excerpt from /tmp/nemoclaw-start.log.
  • Fault tests cover a missing dashboard tail, failed gateway launch, listener-health failure, and crash quarantine.
  • An exact-head live Hermes restart E2E succeeds without rebuilding the sandbox or falling back to a detached/manual gateway launch, with API, dashboard, and forwards healthy afterward.

Related work

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

Labels

area: observabilityLogging, metrics, tracing, diagnostics, or debug outputarea: sandboxOpenShell sandbox lifecycle, runtime, config, or recoveryintegration: hermesHermes integration behaviorneeds: triageAwaiting maintainer classification

Type

No type

Projects

No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions