Skip to content

refactor(hermes): model durable state and restore order explicitly #8009

Description

@jyaunches

Parent epic: #8004

Selected delivery

The first implementation PR for this issue stacks on #8006 and the first proven state-mutation lifecycle slice in #8010.

That PR adds generic staged restore, validation control, publication, and rollback to the shared state engine. Production activation must run only through a provider-bound transaction that excludes every state writer for the complete operation. It must not add Hermes paths, a Hermes durable-state registry, or a gateway-only quiescence callback.

The #7806 replacement follows this foundation and synthesizes the accepted Hermes behavior from #7871 and #7880.

Problem

Hermes state includes static manifest-declared paths and dynamic profile-local state.

Some restore operations have ordering and consistency requirements:

  • referenced scripts must exist before cron definitions become active;
  • related SQLite files must form one consistency unit;
  • agent gateway and scheduler activity must stop before selected state changes; and
  • a failed restore must preserve the prior recoverable state.

A handwritten Hermes registry would duplicate static paths already owned by AgentDefinition. A Hermes branch in the shared state engine would prevent other agents from using the same restore mechanics.

Architecture decision

The shared state engine owns generic restore mechanics:

  • resolve an ordered state plan;
  • stage the selected state;
  • run required validation before publication;
  • preserve the prior live state;
  • publish staged units in plan order;
  • roll back after publication failure;
  • retain recovery state when rollback cannot complete; and
  • report success only after required validation and cleanup complete.

AgentDefinition supplies static manifest-declared state through the contract in #8006.

A bounded Hermes state adapter supplies dynamic Hermes inputs:

  • default-home and named-profile discovery;
  • Hermes consistency-unit selection;
  • Hermes restore-order requirements;
  • cron script-reference and permission validation;
  • SQLite capture and restore rules;
  • messaging-state validation;
  • agent gateway and scheduler quiescence;
  • drain ownership;
  • process-identity binding; and
  • resume behavior.

The shared state engine controls validation timing and commit or rollback. The Hermes adapter supplies Hermes semantic validation.

The shared engine must not contain Hermes paths, commands, profile discovery, drain rules, or restore-order constants.

Do not add a generalized transaction framework. Add only mechanics consumed by the existing state restore path and its tests.

State-mutation security boundary

The shared state engine must not publish into a live agent-owned filesystem through ordinary sandbox-user SSH or exec. A gateway-only drain or a point-in-time process check is not a mutation barrier: gateway, dashboard, scheduler, messaging, terminal, and background processes can share the sandbox UID, and a supervisor can relaunch them.

Before #8009 activates staged restore in production, a provider-bound lifecycle operation must:

  • pin the exact sandbox runtime and lifecycle generation;
  • establish one root-owned, exclusive, nonce-bound transaction;
  • quiesce every process that can mutate the selected state namespace, not only the agent gateway;
  • perform or mediate staged publication and rollback while that boundary remains held;
  • prevent supervisor or concurrent exec relaunch from reopening the namespace;
  • restart services cleanly so no process retains stale file descriptors or in-memory state; and
  • remain fail-closed after identity drift, controller failure, incomplete rollback, or incomplete activation proof.

The operation belongs behind the existing RuntimeProviderBundle and lifecycle authority. Do not add a second provider registry or treat an optional callback as proof of quiescence. Providers that cannot supply the boundary must keep the legacy restore path or reject staged activation explicitly.

Initial generic foundation PR

The stacked PR must:

  • integrate staged publication with the existing restore production path only behind the proven provider-bound state-mutation transaction from refactor(hermes): make NemoClaw the managed gateway lifecycle authority #8010;
  • preserve live state before replacement;
  • publish complete state units in supplied order;
  • run validation before the engine commits the restore;
  • roll back published units after a failure;
  • retain recovery artifacts when rollback or cleanup cannot complete;
  • keep unsupported runtime providers on the legacy path or reject staged activation explicitly;
  • contain no Hermes branches or path constants; and
  • test Hermes and at least one non-Hermes agent.

The PR description must identify the existing direct-restore implementation that it replaces and the exact lifecycle evidence held across publish, rollback, and activation.

#7806 follow-up

Rework #7806 after the generic foundation lands.

The follow-up must:

  • obtain static scripts and cron paths from AgentDefinition;
  • discover supported default-home and named-profile state through the Hermes adapter;
  • stage scripts and cron definitions through the shared state engine;
  • publish scripts before enabling cron definitions;
  • validate every enabled script reference and required permission;
  • quiesce affected Hermes dispatch before publication;
  • preserve an operator-owned drain;
  • release only the drain created by NemoClaw;
  • bind release to the validated Hermes gateway process identity;
  • use the shared rollback and recovery-tree behavior; and
  • preserve the drain when validation, rollback, or activation evidence is incomplete.

#7871 contributes relevant process-identity, named-profile, and fail-closed activation behavior.

#7880 contributes relevant operator-drain, staged-publication, validation, rollback, and recovery-tree behavior.

Do not close #7871 or #7880 until the replacement covers their accepted behavior and regression evidence. Link both PRs from the replacement PR.

Acceptance criteria

  • Static state paths come from AgentDefinition.
  • Dynamic Hermes profiles are discovered without a finite hard-coded profile list.
  • The shared engine contains no Hermes-specific path or lifecycle policy.
  • Restore stages complete state units before publication.
  • Validation runs before the engine commits the restore.
  • A publication failure triggers rollback.
  • A rollback failure preserves the prior recovery state and prevents a later restore from overwriting it.
  • SQLite state uses a documented consistency method, including WAL and SHM handling.
  • Referenced cron scripts validate before cron definitions become active.
  • Messaging metadata has an explicit restore and validation rule when messaging state is in scope.
  • Restore completion means every required validation passed.
  • Snapshots remain portable across supported rebuild and replacement flows.
  • Tests cover stale-base rebuild, replacement restore, partial archive rejection, and rollback.
  • Shared-engine tests cover Hermes and at least one non-Hermes agent.

Delivery order

  1. refactor(hermes): separate configuration, runtime, and durable state ownership #8006 static state authority.
  2. First refactor(hermes): make NemoClaw the managed gateway lifecycle authority #8010 provider-bound state-mutation lifecycle slice.
  3. Generic state-engine staging, validation control, publication, and rollback through that boundary.
  4. [Ubuntu 24.04][Upgrade] rebuild enables restored cron jobs before their scripts are available #7806 Hermes cron and script restore adapter.
  5. SQLite consistency handling.
  6. [Ubuntu 24.04][Upgrade] rebuild does not restore the Slack home channel #7803 messaging-state restore and validation.
  7. Removal of superseded restore branches and path inventories.

Later PRs may stack only when they consume an interface from an earlier PR. A generic engine may be characterized before the lifecycle boundary lands, but it must not be production-activated through sandbox-user SSH or exec.

E2E acceptance

Each PR must follow the PR Review Advisor and E2E contract in #8004.

Minimum live coverage for every behavior-changing PR:

  • rebuild-hermes
  • rebuild-hermes-stale-base
  • state-backup-restore
  • snapshot-commands
  • hermes-e2e
  • gateway-guard-recovery
  • hermes-slack when messaging state is in scope
  • hermes-discord when messaging state is in scope
  • channels-stop-start with the Hermes selector when channel state is in scope

Add an issue-specific live fixture when an existing target does not prove restore ordering. PR Review Advisor can add jobs or targets.

Related acceptance cases

Non-goals

  • Do not add a handwritten Hermes durable-state registry.
  • Do not duplicate static paths outside AgentDefinition.
  • Do not put Hermes profile discovery or lifecycle behavior in the shared state engine.
  • Do not treat process identifiers, listener state, or temporary gateway metadata as durable state without a recovery requirement.
  • Do not report successful restore before required validation completes.
  • Do not add another lifecycle state machine or provider registry.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

Labels

area: architectureArchitecture, design debt, major refactors, or maintainabilityarea: sandboxOpenShell sandbox lifecycle, runtime, config, or recoveryintegration: hermesHermes integration behavior

Type

Projects

No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions