Skip to content

test(e2e): refactor messaging channel coverage - #6053

Closed
sandl99 wants to merge 19 commits into
mainfrom
test/e2e-messaging-refactor-6022
Closed

sandl99 wants to merge 19 commits into
mainfrom
test/e2e-messaging-refactor-6022

Conversation

@sandl99

@sandl99 sandl99 commented Jun 30, 2026 •

Copy link
Copy Markdown
Collaborator

Summary

Refactors the live E2E messaging tests around channel-oriented suites so OpenClaw and Hermes lifecycle coverage share one manifest-aligned helper. Adds Microsoft Teams to the shared lifecycle matrix and adds a hermetic OpenClaw conflict guard target for duplicate channel credentials.

Related Issue

Fixes #6022

Changes

  • Replaces channel-specific add/remove and stop/start live tests with openclaw-channels-* and hermes-channels-* lifecycle suites backed by channels-lifecycle-helpers.ts.
  • Adds Teams lifecycle coverage across OpenClaw and Hermes, including fake CI env, provider/policy/config assertions, and workflow boundary validation.
  • Adds openclaw-channels-conflict-guard to prove duplicate Telegram credentials abort without registry mutation or secret leakage.
  • Renames messaging E2E targets and workflow selectors to channel-scoped names, including pairing, credential rewrite, token rotation, Telegram injection safety, and WhatsApp QR coverage.
  • Updates docs and troubleshooting wording for the current Teams/channel behavior.

Type of Change

  • Code change (feature, bug fix, or refactor)
  • Code change with doc updates
  • Doc only (prose changes, no code sample modifications)
  • Doc only (includes code sample changes)

Quality Gates

  • Tests added or updated for changed behavior
  • Existing tests cover changed behavior — justification:
  • Tests not applicable — justification:
  • Docs updated for user-facing behavior changes
  • Docs not applicable — justification:
  • Sensitive paths changed (security, policy, credentials, preflight, onboarding, inference, runner, sandbox, or messaging)
  • Sensitive-path review completed or maintainer-approved waiver recorded — reviewer/approval link/justification: self-reviewed messaging/test/workflow changes; new conflict guard asserts no secret/hash leakage and no registry mutation.
  • Non-success, skipped, or missing CI check accepted by maintainer — check name, approval link, and follow-up issue:

Verification

  • PR description includes the DCO sign-off declaration and every commit appears as Verified in GitHub
  • Git hooks passed during commit and push, or npx prek run --from-ref main --to-ref HEAD passes
  • Targeted tests pass for changed behavior
  • Full npm test passes (broad runtime changes only)
  • Quality Gates section completed with required justifications or waivers
  • No secrets, API keys, or credentials committed
  • npm run docs builds without warnings (doc changes only)
  • Doc pages follow the style guide (doc changes only)
  • New doc pages include SPDX header and frontmatter (new pages only)

Targeted verification run:

  • npm run build:cli
  • npm run typecheck:cli
  • npx @biomejs/biome lint test/e2e/live/channels-lifecycle-helpers.ts test/e2e/live/openclaw-channels-conflict-guard.test.ts tools/e2e/workflow-boundary.mts test/e2e/support/e2e-workflow.test.ts
  • env NEMOCLAW_RUN_LIVE_E2E=1 npx vitest list --project e2e-live test/e2e/live/openclaw-channels-add-remove.test.ts test/e2e/live/hermes-channels-add-remove.test.ts test/e2e/live/openclaw-channels-stop-start.test.ts test/e2e/live/hermes-channels-stop-start.test.ts test/e2e/live/openclaw-channels-conflict-guard.test.ts
  • env NEMOCLAW_RUN_LIVE_E2E=1 npx vitest run --project e2e-live test/e2e/live/openclaw-channels-conflict-guard.test.ts --silent=false --reporter=default
  • npx vitest run --project e2e-support test/e2e/support/e2e-workflow.test.ts test/e2e/support/openclaw-channels-pairing-workflow-boundary.test.ts --silent=false --reporter=default
  • npx vitest run test/brev-nightly-workflow.test.ts test/e2e-advisor-targets.test.ts test/e2e-release-gate-workflow.test.ts test/regression-e2e-workflow.test.ts --silent=false --reporter=default
  • npm run test:projects:check
  • npm run test:titles:check
  • npm run test:imports:check
  • npm run test-size:check
  • npm run docs passed with 2 pre-existing Fern warnings
  • git diff --check
  • Push hook TypeScript checks passed

Commit pre-hook note: the initial commit hook was not used for the final commit after broad local test-cli and unrelated integration checks failed in this environment; the targeted checks above passed.


Signed-off-by: San Dang sdang@nvidia.com

Summary by CodeRabbit

  • New Features
    • Added new OpenClaw channel-focused live E2E suites (pairing, add/remove, stop/start, token rotation, credential rewrite, conflict guard, Telegram injection safety), plus Hermes channel lifecycle coverage.
    • Updated regression to run the new OpenClaw WhatsApp QR compact pairing job.
  • Bug Fixes
    • Strengthened conflict-guard behavior to prevent duplicate-credential operations from continuing and avoid secret leakage.
    • Refreshed E2E workflow boundary validation to align with the updated channel job set and security checks.
  • Documentation
    • Updated live troubleshooting and sandbox command references to the OpenClaw credential-rewrite lane (including Microsoft Teams).

Signed-off-by: San Dang <sdang@nvidia.com>
@sandl99 sandl99 added the area: docs Documentation, examples, guides, or docs build label Jun 30, 2026
@sandl99 sandl99 self-assigned this Jun 30, 2026
@coderabbitai

coderabbitai Bot commented Jun 30, 2026 •

Copy link
Copy Markdown
Contributor

Review Change Stack

Note

Reviews paused

It looks like this branch is under active development. To avoid overwhelming you with review comments due to an influx of new commits, CodeRabbit has automatically paused this review. You can configure this behavior by changing the reviews.auto_review.auto_pause_after_reviewed_commits setting.

Use the following commands to manage reviews:

  • @coderabbitai resume to resume automatic reviews.
  • @coderabbitai review to trigger a single review.

Use the checkboxes below for quick actions:

  • ▶️ Resume reviews
  • 🔍 Trigger review
📝 Walkthrough

Walkthrough

Renames live E2E coverage to feature-scoped OpenClaw and Hermes channel names, adds shared lifecycle and pairing helpers, introduces conflict-guard coverage, and updates workflow dispatching, boundary validation, regression wiring, docs, and support tests to match the renamed suites and jobs.

Changes

OpenClaw channels E2E refactor

Layer / File(s) Summary
CI workflow job graph updates
.github/workflows/brev-nightly-e2e.yaml, .github/workflows/e2e-branch-validation.yaml, .github/workflows/e2e.yaml, .github/workflows/regression-e2e.yaml
Updates the Brev nightly, branch-validation, main E2E, and regression workflows to use the new OpenClaw channel suite names, renamed jobs, removed Hermes Slack and Discord jobs, and rewritten notification, reporting, and scorecard dependency lists.
Lifecycle tests and workflow validators
test/e2e/live/channels-lifecycle-helpers.ts, test/e2e/live/channels-stop-start-helpers.ts, test/e2e/live/channels-stop-start-safety.ts, test/e2e/live/channels-add-remove.test.ts, test/e2e/live/openclaw-channels-add-remove.test.ts, test/e2e/live/openclaw-channels-stop-start.test.ts, test/e2e/live/hermes-channels-add-remove.test.ts, test/e2e/live/hermes-channels-stop-start.test.ts, tools/e2e/workflow-boundary.mts, test/e2e/support/e2e-workflow.test.ts, test/e2e/support/openclaw-channels-pairing-workflow-boundary.test.ts, test/e2e-release-gate-workflow.test.ts, test/e2e-advisor-targets.test.ts
Introduces shared channel lifecycle helpers, adds OpenClaw and Hermes add-remove and stop-start entrypoints, and rewires workflow-boundary and support tests to validate the renamed lifecycle jobs, pairing boundary job, and OpenClaw-specific stop/start drift rules.
Pairing, credential rewrite, conflict guard, and token rotation
test/e2e/live/openclaw-channels-pairing-discord.ts, test/e2e/live/openclaw-channels-pairing-slack.ts, test/e2e/live/openclaw-channels-pairing-whatsapp-qr.ts, test/e2e/live/openclaw-channels-pairing.test.ts, test/e2e/live/openclaw-channels-credential-rewrite-helpers.ts, test/e2e/live/openclaw-channels-credential-rewrite.test.ts, test/e2e/live/openclaw-channels-conflict-guard.test.ts, test/e2e/live/openclaw-channels-telegram-injection-safety.test.ts, test/e2e/live/openclaw-channels-token-rotation.test.ts, test/e2e/live/openclaw-pairing-helpers.ts, test/e2e/live/phase6-messaging-helpers.ts, test/regression-e2e-workflow.test.ts, test/e2e/live/rebuild-openclaw.test.ts, test/whatsapp-qr-compact.test.ts
Adds the OpenClaw pairing runners for Discord, Slack, and WhatsApp QR compact rendering; introduces the combined pairing test; replaces the messaging-providers credential rewrite suite and renames its artifacts; adds the Telegram conflict-guard test; and retitles the token-rotation and telegram-injection-safety live tests.
Docs and supporting comments
docs/manage-sandboxes/messaging-channels.mdx, docs/reference/troubleshooting.mdx, src/lib/actions/sandbox/policy-explain.ts, test/e2e/brev-e2e.test.ts
Updates Brev E2E suite coverage notes, troubleshooting guidance, policy-explain comments, and related test comments to reference the renamed OpenClaw channel suites and targets.

Estimated code review effort

🎯 4 (Complex) | ⏱️ ~60 minutes

Possibly related PRs

  • NVIDIA/NemoClaw#1092: This PR also updates Brev E2E suite selection and channel-safety coverage naming, intersecting with the same test/e2e/brev-e2e.test.ts branching area.

Suggested labels

refactor

Suggested reviewers

  • cv
  • ericksoa
🚥 Pre-merge checks | ✅ 4 | ❌ 1

❌ Failed checks (1 warning)

Check name Status Explanation Resolution
Docstring Coverage ⚠️ Warning Docstring coverage is 2.44% which is insufficient. The required threshold is 80.00%. Write docstrings for the functions missing them to satisfy the coverage threshold.
✅ Passed checks (4 passed)
Check name Status Explanation
Description Check ✅ Passed Check skipped - CodeRabbit’s high-level summary is enabled.
Title check ✅ Passed The title is concise and accurately describes the main change: refactoring messaging E2E coverage into channel-oriented suites.
Linked Issues check ✅ Passed The PR matches #6022 by renaming and reorganizing lifecycle, pairing, rewrite, conflict, rotation, and injection-safety coverage into channel-based suites.
Out of Scope Changes check ✅ Passed The changes stay within the channel refactor and its workflow, test, and docs support; no clear unrelated scope creep appears.
✨ Finishing Touches
🧪 Generate unit tests (beta)
  • Create PR with unit tests
  • Commit unit tests in branch test/e2e-messaging-refactor-6022

Comment @coderabbitai help to get the list of available commands.

@github-code-quality

github-code-quality Bot commented Jun 30, 2026 •

Copy link
Copy Markdown
Contributor

Code Coverage Overview

Languages: TypeScript

TypeScript / code-coverage/plugin

The overall coverage in the test/e2e-messaging-r... branch is 96%. Coverage data for the main branch is not yet available.

Show a code coverage summary of the most covered files.
File main test/e2e-messaging-r... 630d4d1 +/-
nemoclaw/src/se...cret-scanner.ts — 100% —
nemoclaw/src/commands/slash.ts — 100% —
nemoclaw/src/bl...eprint/state.ts — 98% —
nemoclaw/src/onboard/config.ts — 98% —
nemoclaw/src/bl...int/snapshot.ts — 97% —
nemoclaw/src/blueprint/ssrf.ts — 97% —
nemoclaw/src/bl...print/runner.ts — 95% —
nemoclaw/src/co...ration-state.ts — 94% —
nemoclaw/src/bl...ate-networks.ts — 94% —
nemoclaw/src/index.ts — 94% —

TypeScript / code-coverage/cli

The overall coverage in the test/e2e-messaging-r... branch is 69%. Coverage data for the main branch is not yet available.

Show a code coverage summary of the most covered files.
File main test/e2e-messaging-r... 630d4d1 +/-
src/lib/actions...dbox/rebuild.ts — 82% —
src/lib/actions...all/run-plan.ts — 80% —
src/lib/state/o...oard-session.ts — 79% —
src/lib/shields/index.ts — 75% —
src/lib/state/sandbox.ts — 73% —
src/lib/onboard/preflight.ts — 69% —
src/lib/onboard...er-gpu-patch.ts — 59% —
src/lib/actions...licy-channel.ts — 58% —
src/lib/policy/index.ts — 56% —
src/lib/onboard.ts — 20% —

Updated July 02, 2026 15:43 UTC
Code Coverage is in Public Preview. Learn more and provide us with your feedback.

@github-actions

Copy link
Copy Markdown
Contributor

@github-actions

github-actions Bot commented Jun 30, 2026 •

Copy link
Copy Markdown
Contributor

E2E Advisor Recommendation

Required E2E: network-policy, openclaw-channels-add-remove, hermes-channels-add-remove, openclaw-channels-stop-start, hermes-channels-stop-start, openclaw-channels-credential-rewrite, openclaw-channels-conflict-guard, openclaw-channels-token-rotation, openclaw-channels-telegram-injection-safety, openclaw-channels-pairing, full-e2e
Optional E2E: docs-validation, rebuild-openclaw, launchable-smoke

Dispatch hint: network-policy,openclaw-channels-add-remove,hermes-channels-add-remove,openclaw-channels-stop-start,hermes-channels-stop-start,openclaw-channels-credential-rewrite,openclaw-channels-conflict-guard,openclaw-channels-token-rotation,openclaw-channels-telegram-injection-safety,openclaw-channels-pairing,full-e2e

Workflow run

Full advisor summary

E2E Recommendation Advisor

Base: origin/main
Head: HEAD
Confidence: medium

Required E2E

  • network-policy (high): Validates the real OpenShell sandbox boundary for policy-add/policy-remove, in-sandbox policy context refresh, and network policy allow/deny behavior after the sandbox policy write path changed.
  • openclaw-channels-add-remove (high): Required because OpenClaw channel lifecycle tests/workflow changed and this job exercises no-channel onboarding, channel add/remove, rebuild, credential reuse, policy-list, and rendered config cleanup across supported channels.
  • hermes-channels-add-remove (high): Required because Hermes channel lifecycle coverage was added/rewired and must validate add/remove, rebuild, credentials, policy-list, and per-agent rendered config behavior in a real sandbox.
  • openclaw-channels-stop-start (high): Required because stop/start channel workflow and helpers changed; validates OpenClaw messaging channel stop/start, rebuild, provider reuse, registry, policy-list, and in-sandbox config contracts.
  • hermes-channels-stop-start (high): Required because the former matrixed stop/start path was split and Hermes now has a discrete job; validates Hermes channel stop/start, rebuild, credential/provider reuse, and in-sandbox config contracts.
  • openclaw-channels-credential-rewrite (high): Required because credential rewrite/provider coverage and branch/nightly suite names changed; validates provider creation, credential isolation, config patching, reachability, token rewriting, and secret-boundary behavior.
  • openclaw-channels-conflict-guard (medium): Required because the PR introduces/rewires the duplicate credential conflict guard job; validates the credential-hash conflict boundary for OpenClaw channels in a real workflow job.
  • openclaw-channels-token-rotation (medium): Required because token rotation job/test names changed; validates fake Telegram/Discord/Slack credential rotation boundaries and ensures the renamed workflow selector still dispatches correctly.
  • openclaw-channels-telegram-injection-safety (medium): Required for security coverage because the Telegram injection job was renamed and rewired; validates shell metacharacter injection prevention, process-table leak checks, and SANDBOX_NAME validation across the real sandbox boundary.
  • openclaw-channels-pairing (high): Required because Discord/Slack/WhatsApp pairing jobs were consolidated; validates Discord Gateway and Slack Socket Mode connect-shell approval plus WhatsApp QR rendering under the new workflow selector.
  • full-e2e (high): Required as an end-to-end OpenClaw user-flow smoke because the workflow dependencies were changed from legacy channel job names to the new split channel jobs, and this verifies install/onboard/sandbox/hosted inference still works through the updated graph.

Optional E2E

  • docs-validation (low): Useful for the changed messaging-channel and troubleshooting MDX docs; not merge-blocking for runtime because the docs changes do not alter sandbox behavior.
  • rebuild-openclaw (high): Optional adjacent confidence because rebuild-openclaw live coverage changed and channel add/remove relies on rebuild behavior, but the channel jobs above already exercise the critical rebuild path for this PR.
  • launchable-smoke (medium): Optional confidence for the changed Brev/nightly/branch-validation workflow wiring and launchable-facing comments; useful if reviewers want a quick public install/launchable smoke beyond the channel-specific jobs.

New E2E recommendations

  • e2e-workflow-orchestration (medium): The PR renames branch-validation suites and many workflow selectors, but there is no cheap live GitHub Actions dispatch smoke that validates renamed selector inputs across e2e.yaml, e2e-branch-validation.yaml, brev-nightly-e2e.yaml, and regression-e2e.yaml without provisioning full Brev resources.
    • Suggested test: workflow-dispatch-selector-smoke

Dispatch hint

  • Workflow: e2e.yaml
  • jobs input: network-policy,openclaw-channels-add-remove,hermes-channels-add-remove,openclaw-channels-stop-start,hermes-channels-stop-start,openclaw-channels-credential-rewrite,openclaw-channels-conflict-guard,openclaw-channels-token-rotation,openclaw-channels-telegram-injection-safety,openclaw-channels-pairing,full-e2e

@github-actions

github-actions Bot commented Jun 30, 2026 •

Copy link
Copy Markdown
Contributor

E2E Target Recommendation

Required E2E targets: hermes-channels-add-remove, hermes-channels-stop-start, openclaw-channels-add-remove, openclaw-channels-conflict-guard, openclaw-channels-credential-rewrite, openclaw-channels-pairing, openclaw-channels-stop-start, openclaw-channels-telegram-injection-safety, openclaw-channels-token-rotation, rebuild-openclaw, e2e-all
Optional E2E targets: None

Dispatch required E2E targets:

  • gh workflow run e2e.yaml --ref <pr-head-ref> --field jobs=hermes-channels-add-remove
  • gh workflow run e2e.yaml --ref <pr-head-ref> --field jobs=hermes-channels-stop-start
  • gh workflow run e2e.yaml --ref <pr-head-ref> --field jobs=openclaw-channels-add-remove
  • gh workflow run e2e.yaml --ref <pr-head-ref> --field jobs=openclaw-channels-conflict-guard
  • gh workflow run e2e.yaml --ref <pr-head-ref> --field jobs=openclaw-channels-credential-rewrite
  • gh workflow run e2e.yaml --ref <pr-head-ref> --field jobs=openclaw-channels-pairing
  • gh workflow run e2e.yaml --ref <pr-head-ref> --field jobs=openclaw-channels-stop-start
  • gh workflow run e2e.yaml --ref <pr-head-ref> --field jobs=openclaw-channels-telegram-injection-safety
  • gh workflow run e2e.yaml --ref <pr-head-ref> --field jobs=openclaw-channels-token-rotation
  • gh workflow run e2e.yaml --ref <pr-head-ref> --field jobs=rebuild-openclaw
  • gh workflow run e2e.yaml --ref <pr-head-ref>

Workflow run

Full E2E target advisor summary

E2E Target Advisor

Base: origin/main
Head: HEAD
Confidence: high

Required E2E targets

  • hermes-channels-add-remove: Focused free-standing E2E job wired for changed live test test/e2e/live/hermes-channels-add-remove.test.ts.
    • Dispatch: gh workflow run e2e.yaml --ref <pr-head-ref> --field jobs=hermes-channels-add-remove
  • hermes-channels-stop-start: Focused free-standing E2E job wired for changed live test test/e2e/live/hermes-channels-stop-start.test.ts.
    • Dispatch: gh workflow run e2e.yaml --ref <pr-head-ref> --field jobs=hermes-channels-stop-start
  • openclaw-channels-add-remove: Focused free-standing E2E job wired for changed live test test/e2e/live/openclaw-channels-add-remove.test.ts.
    • Dispatch: gh workflow run e2e.yaml --ref <pr-head-ref> --field jobs=openclaw-channels-add-remove
  • openclaw-channels-conflict-guard: Focused free-standing E2E job wired for changed live test test/e2e/live/openclaw-channels-conflict-guard.test.ts.
    • Dispatch: gh workflow run e2e.yaml --ref <pr-head-ref> --field jobs=openclaw-channels-conflict-guard
  • openclaw-channels-credential-rewrite: Focused free-standing E2E job wired for changed live test test/e2e/live/openclaw-channels-credential-rewrite.test.ts.
    • Dispatch: gh workflow run e2e.yaml --ref <pr-head-ref> --field jobs=openclaw-channels-credential-rewrite
  • openclaw-channels-pairing: Focused free-standing E2E job wired for changed live test test/e2e/live/openclaw-channels-pairing.test.ts.
    • Dispatch: gh workflow run e2e.yaml --ref <pr-head-ref> --field jobs=openclaw-channels-pairing
  • openclaw-channels-stop-start: Focused free-standing E2E job wired for changed live test test/e2e/live/openclaw-channels-stop-start.test.ts.
    • Dispatch: gh workflow run e2e.yaml --ref <pr-head-ref> --field jobs=openclaw-channels-stop-start
  • openclaw-channels-telegram-injection-safety: Focused free-standing E2E job wired for changed live test test/e2e/live/openclaw-channels-telegram-injection-safety.test.ts.
    • Dispatch: gh workflow run e2e.yaml --ref <pr-head-ref> --field jobs=openclaw-channels-telegram-injection-safety
  • openclaw-channels-token-rotation: Focused free-standing E2E job wired for changed live test test/e2e/live/openclaw-channels-token-rotation.test.ts.
    • Dispatch: gh workflow run e2e.yaml --ref <pr-head-ref> --field jobs=openclaw-channels-token-rotation
  • rebuild-openclaw: Focused free-standing E2E job wired for changed live test test/e2e/live/rebuild-openclaw.test.ts.
    • Dispatch: gh workflow run e2e.yaml --ref <pr-head-ref> --field jobs=rebuild-openclaw
  • e2e-all: The PR changes the canonical E2E workflow machinery and free-standing live job wiring in .github/workflows/e2e.yaml, including renamed/split messaging channel jobs, dependency/reporting wiring, and related workflow-boundary/support tests. A full e2e.yaml fan-out is required to validate the shared dispatch/matrix/default-job behavior and all affected live E2E target paths.
    • Dispatch: gh workflow run e2e.yaml --ref <pr-head-ref>

Optional E2E targets

  • None.

Relevant changed files

  • .github/workflows/e2e.yaml
  • test/e2e/brev-e2e.test.ts
  • test/e2e/live/channels-add-remove.test.ts
  • test/e2e/live/channels-lifecycle-helpers.ts
  • test/e2e/live/channels-stop-start-helpers.ts
  • test/e2e/live/channels-stop-start-safety.ts
  • test/e2e/live/channels-stop-start.test.ts
  • test/e2e/live/hermes-channels-add-remove.test.ts
  • test/e2e/live/hermes-channels-stop-start.test.ts
  • test/e2e/live/hermes-discord.test.ts
  • test/e2e/live/hermes-slack-e2e-helpers.ts
  • test/e2e/live/hermes-slack-e2e.test.ts
  • test/e2e/live/openclaw-channels-add-remove.test.ts
  • test/e2e/live/openclaw-channels-conflict-guard.test.ts
  • test/e2e/live/openclaw-channels-credential-rewrite-helpers.ts
  • test/e2e/live/openclaw-channels-credential-rewrite.test.ts
  • test/e2e/live/openclaw-channels-pairing-discord.ts
  • test/e2e/live/openclaw-channels-pairing-slack.ts
  • test/e2e/live/openclaw-channels-pairing-whatsapp-qr.ts
  • test/e2e/live/openclaw-channels-pairing.test.ts
  • test/e2e/live/openclaw-channels-stop-start.test.ts
  • test/e2e/live/openclaw-channels-telegram-injection-safety.test.ts
  • test/e2e/live/openclaw-channels-token-rotation.test.ts
  • test/e2e/live/openclaw-discord-pairing.test.ts
  • test/e2e/live/openclaw-pairing-helpers.ts
  • test/e2e/live/openclaw-slack-pairing.test.ts
  • test/e2e/live/phase6-messaging-helpers.ts
  • test/e2e/live/rebuild-openclaw.test.ts
  • test/e2e/support/channels-lifecycle-workflow-boundary.test.ts
  • test/e2e/support/dockerhub-auth-workflow-boundary.test.ts
  • test/e2e/support/e2e-workflow.test.ts
  • test/e2e/support/openclaw-channels-pairing-workflow-boundary.test.ts
  • test/e2e/support/openclaw-discord-workflow-boundary.test.ts
  • tools/e2e/upload-e2e-artifacts-workflow-boundary.mts
  • tools/e2e/workflow-boundary.mts

@github-actions

github-actions Bot commented Jun 30, 2026 •

Copy link
Copy Markdown
Contributor

PR Review Advisor — No blocking findings

Merge posture: No blocking advisor findings
Primary next action: Add or justify PRA-T1 and any related test follow-ups.
Open items: 0 required · 0 warnings · 0 suggestions · 4 test follow-ups
Since last review: 0 prior items resolved · 0 still apply · 0 new items found

Action checklist

  • PRA-T1 Add or justify test follow-up: Runtime validation
  • PRA-T2 Add or justify test follow-up: Runtime validation
  • PRA-T3 Add or justify test follow-up: Runtime validation
  • PRA-T4 Add or justify test follow-up: Runtime validation
Test follow-ups to resolve or justify

If these cover changed behavior, prefer adding them in this PR; otherwise state why existing coverage is enough or link the follow-up.

  • PRA-T1 Runtime validation — Verify openclaw-channels-add-remove and hermes-channels-add-remove exercise every channel and leave no raw channel tokens in sandbox config after rebuild.. The changed behavior is primarily workflow and live Docker/OpenShell/channel lifecycle coverage. Static boundary tests are extensive, but the confidence-critical behavior depends on real CLI, sandbox, provider, policy, and artifact boundaries.
  • PRA-T2 Runtime validation — Verify openclaw-channels-stop-start and hermes-channels-stop-start preserve providers while toggling policy/config state through stop, rebuild, start, and rebuild.. The changed behavior is primarily workflow and live Docker/OpenShell/channel lifecycle coverage. Static boundary tests are extensive, but the confidence-critical behavior depends on real CLI, sandbox, provider, policy, and artifact boundaries.
  • PRA-T3 Runtime validation — Verify openclaw-channels-conflict-guard aborts duplicate credential adds without registry mutation or token/hash disclosure for each credentialed channel.. The changed behavior is primarily workflow and live Docker/OpenShell/channel lifecycle coverage. Static boundary tests are extensive, but the confidence-critical behavior depends on real CLI, sandbox, provider, policy, and artifact boundaries.
  • PRA-T4 Runtime validation — Verify workflow selector validation accepts the renamed channel jobs and rejects malformed jobs/targets before secret-bearing E2E jobs run.. The changed behavior is primarily workflow and live Docker/OpenShell/channel lifecycle coverage. Static boundary tests are extensive, but the confidence-critical behavior depends on real CLI, sandbox, provider, policy, and artifact boundaries.

Workflow run details

This is an automated, non-binding review; it still expects maintainers and agents to respond to each required or warning item. Treat suggestions as current-PR improvements when they touch changed code; defer only with maintainer rationale or a linked follow-up. A human maintainer must make the final merge decision.

@github-actions

github-actions Bot commented Jun 30, 2026 •

Copy link
Copy Markdown
Contributor

PR Review Advisor (Nemotron Ultra) — Blocked

Merge posture: Do not merge until addressed
Primary next action: Fix PRA-6: workflow-boundary.mts trusted validator lacks self-test coverage; then add or justify PRA-T1.
Open items: 8 required · 13 warnings · 3 suggestions · 8 test follow-ups
Since last review: 0 prior items resolved · 15 still apply · 5 new items found

Action checklist

  • PRA-6 Fix: workflow-boundary.mts trusted validator lacks self-test coverage in tools/e2e/workflow-boundary.mts:1
  • PRA-7 Fix: Missing validateHermesChannelsAddRemove and validateHermesChannelsStopStart validators in tools/e2e/workflow-boundary.mts:1
  • PRA-8 Fix: No Hermes workflow boundary validator tests exist in test/e2e/support:1
  • PRA-9 Fix: Missing Hermes equivalent of credential rewrite test in test/e2e/live:1
  • PRA-10 Fix: Missing Hermes pairing tests for Discord Gateway and Slack Socket Mode in test/e2e/live:1
  • PRA-11 Fix: Missing hermes-channels-conflict-guard workflow job in .github/workflows/e2e.yaml:3978
  • PRA-12 Fix: Conflict guard test not parameterized over AgentKind in test/e2e/live/openclaw-channels-conflict-guard.test.ts:202
  • PRA-19 Fix: Missing validators for openclaw-channels-add-remove and openclaw-channels-stop-start jobs in tools/e2e/workflow-boundary.mts:1
  • PRA-1 Resolve or justify: Source-of-truth review needed: test/e2e/live/channels-lifecycle-helpers.ts:359
  • PRA-2 Resolve or justify: Source-of-truth review needed: src/lib/agent/onboard.ts:1
  • PRA-3 Resolve or justify: Source-of-truth review needed: src/lib/build-context.ts:27
  • PRA-4 Resolve or justify: Source-of-truth review needed: tools/e2e/workflow-boundary.mts:1
  • PRA-5 Resolve or justify: Source-of-truth review needed: test/e2e/live/issue-2478-crash-loop-recovery.test.ts:1
  • PRA-13 Resolve or justify: workflow-boundary.mts lacks source-of-truth documentation in tools/e2e/workflow-boundary.mts:1
  • PRA-14 Resolve or justify: channels-lifecycle-helpers.ts is a 1059-line monolith without architecture justification in test/e2e/live/channels-lifecycle-helpers.ts:1
  • PRA-15 Resolve or justify: hermesChannelProbe workaround lacks documentation and removal condition in test/e2e/live/channels-lifecycle-helpers.ts:359
  • PRA-16 Resolve or justify: src/lib/agent/onboard.ts is a 637-line monolith with no helper extraction in src/lib/agent/onboard.ts:1
  • PRA-17 Resolve or justify: Build context exclusion patterns duplicated across 4 files with imperfect synchronization in src/lib/build-context.ts:27
  • PRA-T1 Add or justify test follow-up: Runtime validation
  • PRA-T2 Add or justify test follow-up: Runtime validation
  • PRA-T3 Add or justify test follow-up: Runtime validation
  • PRA-T4 Add or justify test follow-up: Runtime validation
  • PRA-T5 Add or justify test follow-up: Runtime validation
  • PRA-T6 Add or justify test follow-up: workflow-boundary.mts trusted validator lacks self-test coverage
  • PRA-T7 Add or justify test follow-up: No Hermes workflow boundary validator tests exist
  • PRA-T8 Add or justify test follow-up: Missing hermes-channels-conflict-guard workflow job
  • PRA-22 In-scope improvement: ensureAgentBaseImage function can be extracted to agent-base-image.ts helper in src/lib/agent/onboard.ts:150
  • PRA-23 In-scope improvement: createAgentSandbox function can be extracted to agent-sandbox-factory.ts helper in src/lib/agent/onboard.ts:160
  • PRA-24 In-scope improvement: Dashboard printing functions can be extracted to agent-dashboard.ts helper in src/lib/agent/onboard.ts:550

Findings index

ID Severity Category Location Required action
PRA-1 Resolve/justify architecture — Identify the invalid state, source boundary, source-fix constraint, regression test, and removal condition before merging the localized behavior.
PRA-2 Resolve/justify architecture — Identify the invalid state, source boundary, source-fix constraint, regression test, and removal condition before merging the localized behavior.
PRA-3 Resolve/justify architecture — Identify the invalid state, source boundary, source-fix constraint, regression test, and removal condition before merging the localized behavior.
PRA-4 Resolve/justify architecture — Identify the invalid state, source boundary, source-fix constraint, regression test, and removal condition before merging the localized behavior.
PRA-5 Resolve/justify architecture — Identify the invalid state, source boundary, source-fix constraint, regression test, and removal condition before merging the localized behavior.
PRA-6 Required tests tools/e2e/workflow-boundary.mts:1 Add test/e2e/support/workflow-boundary-validator.test.ts with fixture-based tests covering: validateE2eWorkflowBoundary (valid/invalid workflow YAML), evaluateE2eWorkflowDispatchSelectors (selector parsing, unknown job rejection), validateFreeStandingWorkflowInventory (duplicate detection, missing keys), and readFreeStandingJobsInventory (caching, error propagation).
PRA-7 Required security tools/e2e/workflow-boundary.mts:1 Add validateHermesChannelsAddRemove and validateHermesChannelsStopStart functions mirroring the OpenClaw validators, checking all security properties. Call them from validateE2eWorkflowBoundary.
PRA-8 Required tests test/e2e/support:1 Add test/e2e/support/hermes-channels-add-remove-workflow-boundary.test.ts and test/e2e/support/hermes-channels-stop-start-workflow-boundary.test.ts mirroring the OpenClaw pattern.
PRA-9 Required security test/e2e/live:1 Add hermes-channels-credential-rewrite.test.ts covering Hermes provider placeholder verification (/sandbox/.hermes/.env), L7 proxy rewrite proofs for Discord Gateway (WebSocket), Slack REST (auth.test, apps.connections.open), Slack Socket Mode (HELLO/EVENT/ACK), and raw token leak checks for Hermes config paths.
PRA-10 Required correctness test/e2e/live:1 Add hermes-channels-pairing-discord.ts and hermes-channels-pairing-slack.ts mirroring OpenClaw patterns. Add corresponding workflow jobs in e2e.yaml.
PRA-11 Required tests .github/workflows/e2e.yaml:3978 Add hermes-channels-conflict-guard job to e2e.yaml mirroring openclaw-channels-conflict-guard. Add to report-to-pr needs list.
PRA-12 Required security test/e2e/live/openclaw-channels-conflict-guard.test.ts:202 Parameterize the test loop over both AgentKind and CHANNELS. Update CHANNEL_FIXTURES to include agent-specific placeholder formats: Hermes WeChat uses WEIXIN_TOKEN, OpenClaw uses openclaw-weixin; Hermes Discord/Slack/Telegram/Teams use same keys but different config paths. Loop over ["openclaw", "hermes"] × CREDENTIALED_CHANNELS.
PRA-13 Resolve/justify architecture tools/e2e/workflow-boundary.mts:1 Add header comment to workflow-boundary.mts documenting it as the source of truth for workflow security properties. Reference the self-test file from PRA-2 once created.
PRA-14 Resolve/justify architecture test/e2e/live/channels-lifecycle-helpers.ts:1 Either split into channels-registry-helpers.ts, channels-provider-helpers.ts, channels-policy-helpers.ts, channels-config-helpers.ts, channels-protocol-helpers.ts if shared assertions are stable, OR add a header comment explaining why the monolith is correct (e.g., frequent cross-lifecycle refactoring, tightly coupled state machine).
PRA-15 Resolve/justify architecture test/e2e/live/channels-lifecycle-helpers.ts:359 Document hermesChannelProbe as a known workaround with explicit removal condition in code comments. Reference expectHermesProtocolCredentialRewrite as the regression guard that will replace it. Call expectHermesProtocolCredentialRewrite from runChannelsAddRemoveTarget for Hermes agent as well.
PRA-16 Resolve/justify architecture src/lib/agent/onboard.ts:1 Extract cohesive helpers before merge: agent-base-image.ts (ensureAgentBaseImage), agent-sandbox-factory.ts (createAgentSandbox), agent-dashboard.ts (getAgentDashboardInfo, printDashboardUi, printAdditionalForwardPorts, resolveUrlPort). Update onboard.ts to import from them.
PRA-17 Resolve/justify architecture src/lib/build-context.ts:27 Centralize exclusion patterns in a single source file (e.g., src/lib/core/build-context-exclusions.ts) that both build-context.ts and a sync script for .dockerignore/.gitignore can reference. At minimum, add a comment in build-context.ts referencing .dockerignore and .gitignore to keep them in sync.
PRA-18 Resolve/justify workflow .github/workflows/e2e.yaml:3342 Add hermes-channels-conflict-guard job to e2e.yaml mirroring openclaw-channels-conflict-guard. Add to report-to-pr needs list.
PRA-19 Required security tools/e2e/workflow-boundary.mts:1 Add validateOpenClawChannelsAddRemoveJob and validateOpenClawChannelsStopStartJob functions mirroring validateOpenClawChannelsCredentialRewriteJob pattern. Call them from validateE2eWorkflowBoundary.
PRA-20 Resolve/justify tests test/e2e/support/channels-lifecycle-workflow-boundary.test.ts:1 Once PRA-18 and PRA-3 are resolved (job-specific validators added), update channels-lifecycle-workflow-boundary.test.ts to test the new validators, or add dedicated test files for each job.

🚨 Required before merge

Address these before merging unless a maintainer explicitly overrides the advisor with rationale.

PRA-6 Required — workflow-boundary.mts trusted validator lacks self-test coverage

  • Location: tools/e2e/workflow-boundary.mts:1
  • Category: tests
  • Problem: The workflow-boundary.mts validator enforces security properties (action SHA pinning, secret non-exposure, checkout persist-credentials=false, artifact path hygiene, DOCKER_CONFIG handling) for all free-standing E2E jobs but has no tests validating its own logic. A bug in the validator could silently allow unsafe workflow patterns.
  • Impact: Trusted code boundary without regression tests; validator drift could permit unpinned actions, secret leakage, or unsafe artifact paths in production workflows.
  • Required action: Add test/e2e/support/workflow-boundary-validator.test.ts with fixture-based tests covering: validateE2eWorkflowBoundary (valid/invalid workflow YAML), evaluateE2eWorkflowDispatchSelectors (selector parsing, unknown job rejection), validateFreeStandingWorkflowInventory (duplicate detection, missing keys), and readFreeStandingJobsInventory (caching, error propagation).
  • Expected follow-up: Fix before merge or get explicit maintainer override.
  • Verification: Check if test/e2e/support/workflow-boundary-validator.test.ts exists and covers the four validator functions with valid/invalid fixtures.
  • Missing regression test: test/e2e/support/workflow-boundary-validator.test.ts with fixture-based tests for all four exported validator functions
  • Done when: The required change is committed and verification passes: Check if test/e2e/support/workflow-boundary-validator.test.ts exists and covers the four validator functions with valid/invalid fixtures.
  • Evidence: tools/e2e/workflow-boundary.mts exports validateE2eWorkflowBoundary, evaluateE2eWorkflowDispatchSelectors, validateFreeStandingWorkflowInventory, readFreeStandingJobsInventory but no test file exercises them

PRA-7 Required — Missing validateHermesChannelsAddRemove and validateHermesChannelsStopStart validators

  • Location: tools/e2e/workflow-boundary.mts:1
  • Category: security
  • Problem: The hermes-channels-add-remove and hermes-channels-stop-start jobs exist in e2e.yaml but have no dedicated validator functions checking: hosted inference config, fake tokens, Docker auth isolation (RUNNER_TEMP), test file path, artifact upload path, checkout persist-credentials=false, action SHA pinning.
  • Impact: Hermes channels jobs bypass job-specific security checks. Only generic inventory validation applies, leaving fake token enforcement, Docker auth isolation, checkout safety, and action pinning unchecked for these jobs.
  • Required action: Add validateHermesChannelsAddRemove and validateHermesChannelsStopStart functions mirroring the OpenClaw validators, checking all security properties. Call them from validateE2eWorkflowBoundary.
  • Expected follow-up: Fix before merge or get explicit maintainer override.
  • Verification: grep for 'validateHermesChannelsAddRemove' and 'validateHermesChannelsStopStart' in tools/e2e/workflow-boundary.mts; verify they are called in validateE2eWorkflowBoundary.
  • Missing regression test: test/e2e/support/hermes-channels-add-remove-workflow-boundary.test.ts and hermes-channels-stop-start-workflow-boundary.test.ts mirroring openclaw-channels-pairing-workflow-boundary.test.ts pattern
  • Done when: The required change is committed and verification passes: grep for 'validateHermesChannelsAddRemove' and 'validateHermesChannelsStopStart' in tools/e2e/workflow-boundary.mts; verify they are called in validateE2eWorkflowBoundary.
  • Evidence: tools/e2e/workflow-boundary.mts only has validators for openclaw-channels-credential-rewrite, openclaw-channels-conflict-guard, openclaw-channels-pairing, openclaw-channels-telegram-injection-safety — no Hermes channels add-remove/stop-start validators

PRA-8 Required — No Hermes workflow boundary validator tests exist

  • Location: test/e2e/support:1
  • Category: tests
  • Problem: The OpenClaw pattern has test/e2e/support/openclaw-channels-pairing-workflow-boundary.test.ts that mutates the workflow and asserts validateE2eWorkflowBoundary catches security violations. No equivalent exists for Hermes channels jobs.
  • Impact: Workflow boundary drift for Hermes channels jobs would not be caught by tests. Security properties like fake token enforcement, Docker auth isolation, and action pinning have no regression test coverage.
  • Required action: Add test/e2e/support/hermes-channels-add-remove-workflow-boundary.test.ts and test/e2e/support/hermes-channels-stop-start-workflow-boundary.test.ts mirroring the OpenClaw pattern.
  • Expected follow-up: Fix before merge or get explicit maintainer override.
  • Verification: Check if test/e2e/support/hermes-channels-add-remove-workflow-boundary.test.ts and hermes-channels-stop-start-workflow-boundary.test.ts exist.
  • Missing regression test: test/e2e/support/hermes-channels-add-remove-workflow-boundary.test.ts and hermes-channels-stop-start-workflow-boundary.test.ts
  • Done when: The required change is committed and verification passes: Check if test/e2e/support/hermes-channels-add-remove-workflow-boundary.test.ts and hermes-channels-stop-start-workflow-boundary.test.ts exist.
  • Evidence: Only test/e2e/support/hermes-workflow-boundary.test.ts exists, which only tests hermes-e2e job model pinning, not channels add-remove/stop-start jobs

PRA-9 Required — Missing Hermes equivalent of credential rewrite test

  • Location: test/e2e/live:1
  • Category: security
  • Problem: OpenClaw has test/e2e/live/openclaw-channels-credential-rewrite.test.ts covering: provider placeholder verification (/sandbox/.hermes/.env), L7 proxy rewrite proofs for Discord Gateway (WebSocket), Slack REST (auth.test, apps.connections.open), Slack Socket Mode (HELLO/EVENT/ACK), and raw token leak checks for Hermes config paths. No Hermes version exists.
  • Impact: Hermes agent credential rewrite path (L7 proxy token rewriting for Discord/Slack) has no E2E coverage. Token leakage or rewrite failures in Hermes config paths (/sandbox/.hermes/.env) would go undetected.
  • Required action: Add hermes-channels-credential-rewrite.test.ts covering Hermes provider placeholder verification (/sandbox/.hermes/.env), L7 proxy rewrite proofs for Discord Gateway (WebSocket), Slack REST (auth.test, apps.connections.open), Slack Socket Mode (HELLO/EVENT/ACK), and raw token leak checks for Hermes config paths.
  • Expected follow-up: Fix before merge or get explicit maintainer override.
  • Verification: Check if test/e2e/live/hermes-channels-credential-rewrite.test.ts exists with Discord Gateway WebSocket, Slack REST, and Slack Socket Mode rewrite proofs for Hermes.
  • Missing regression test: test/e2e/live/hermes-channels-credential-rewrite.test.ts with Discord Gateway, Slack REST, Slack Socket Mode, and raw token leak proofs for Hermes
  • Done when: The required change is committed and verification passes: Check if test/e2e/live/hermes-channels-credential-rewrite.test.ts exists with Discord Gateway WebSocket, Slack REST, and Slack Socket Mode rewrite proofs for Hermes.
  • Evidence: No test/e2e/live/hermes-channels-credential-rewrite.test.ts file exists; only openclaw-channels-credential-rewrite.test.ts

PRA-10 Required — Missing Hermes pairing tests for Discord Gateway and Slack Socket Mode

  • Location: test/e2e/live:1
  • Category: correctness
  • Problem: OpenClaw has openclaw-channels-pairing-discord.ts and openclaw-channels-pairing-slack.ts using startFakeDiscordGateway/startFakeSlackApi, applyFakePolicy with websocket rewrite, runDiscordGatewayProof/runSlackSocketModeProof with capture assertions, issuePairingRequest, approveAndAssertPairing. No Hermes equivalents exist.
  • Impact: Hermes agent Discord/Slack pairing flows (connect-shell approval, WebSocket/Socket Mode credential rewrite) have no E2E coverage. Pairing failures or token leakage in Hermes paths would go undetected.
  • Required action: Add hermes-channels-pairing-discord.ts and hermes-channels-pairing-slack.ts mirroring OpenClaw patterns. Add corresponding workflow jobs in e2e.yaml.
  • Expected follow-up: Fix before merge or get explicit maintainer override.
  • Verification: Check if test/e2e/live/hermes-channels-pairing-discord.ts and hermes-channels-pairing-slack.ts exist with fake gateway/API servers and rewrite proofs. Verify corresponding jobs in e2e.yaml.
  • Missing regression test: test/e2e/live/hermes-channels-pairing-discord.ts and hermes-channels-pairing-slack.ts with corresponding e2e.yaml jobs
  • Done when: The required change is committed and verification passes: Check if test/e2e/live/hermes-channels-pairing-discord.ts and hermes-channels-pairing-slack.ts exist with fake gateway/API servers and rewrite proofs. Verify corresponding jobs in e2e.yaml.
  • Evidence: No test/e2e/live/hermes-channels-pairing-discord.ts or hermes-channels-pairing-slack.ts files exist; only openclaw-channels-pairing-*

PRA-11 Required — Missing hermes-channels-conflict-guard workflow job

  • Location: .github/workflows/e2e.yaml:3978
  • Category: tests
  • Problem: OpenClaw has openclaw-channels-conflict-guard job in e2e.yaml but no Hermes equivalent.
  • Impact: Hermes agent duplicate credential conflict detection (preventing two sandboxes using the same bot token) has no CI coverage. Credential conflict bugs in Hermes path would go undetected.
  • Required action: Add hermes-channels-conflict-guard job to e2e.yaml mirroring openclaw-channels-conflict-guard. Add to report-to-pr needs list.
  • Expected follow-up: Fix before merge or get explicit maintainer override.
  • Verification: grep for 'hermes-channels-conflict-guard' in .github/workflows/e2e.yaml; verify it's in the report-to-pr needs array.
  • Missing regression test: hermes-channels-conflict-guard job in e2e.yaml with report-to-pr dependency
  • Done when: The required change is committed and verification passes: grep for 'hermes-channels-conflict-guard' in .github/workflows/e2e.yaml; verify it's in the report-to-pr needs array.
  • Evidence: grep shows openclaw-channels-conflict-guard job at line 3339 but no hermes-channels-conflict-guard

PRA-12 Required — Conflict guard test not parameterized over AgentKind

  • Location: test/e2e/live/openclaw-channels-conflict-guard.test.ts:202
  • Category: security
  • Problem: The test only tests 'openclaw' agent (hardcoded in channelPlanWithCredential agent field). Hermes WeChat uses WEIXIN_TOKEN placeholder vs openclaw-weixin for OpenClaw; Hermes Discord/Slack/Telegram/Teams use same keys but different config paths (/sandbox/.hermes/.env vs /sandbox/.openclaw/openclaw.json).
  • Impact: Conflict guard behavior for Hermes agent (WeChat placeholder format, config path differences) is untested. A Hermes-specific conflict guard bug would not be caught.
  • Required action: Parameterize the test loop over both AgentKind and CHANNELS. Update CHANNEL_FIXTURES to include agent-specific placeholder formats: Hermes WeChat uses WEIXIN_TOKEN, OpenClaw uses openclaw-weixin; Hermes Discord/Slack/Telegram/Teams use same keys but different config paths. Loop over ["openclaw", "hermes"] × CREDENTIALED_CHANNELS.
  • Expected follow-up: Fix before merge or get explicit maintainer override.
  • Verification: Check if openclaw-channels-conflict-guard.test.ts loops over AgentKind and uses agent-specific placeholders/config paths.
  • Missing regression test: Parameterized conflict guard test covering both agents × all credentialed channels with agent-specific placeholders
  • Done when: The required change is committed and verification passes: Check if openclaw-channels-conflict-guard.test.ts loops over AgentKind and uses agent-specific placeholders/config paths.
  • Evidence: channelPlanWithCredential hardcodes agent: "openclaw" at line 194; CHANNEL_FIXTURES only defines openclaw placeholders (e.g., wechat placeholder: "openshell:resolve:env:WECHAT_BOT_TOKEN" but Hermes uses WEIXIN_TOKEN)

PRA-19 Required — Missing validators for openclaw-channels-add-remove and openclaw-channels-stop-start jobs

  • Location: tools/e2e/workflow-boundary.mts:1
  • Category: security
  • Problem: These jobs exist in e2e.yaml but have no dedicated validateOpenClawChannelsAddRemoveJob or validateOpenClawChannelsStopStartJob functions. Only generic inventory validation applies.
  • Impact: OpenClaw channels add/remove and stop/start jobs only get generic inventory validation. Job-specific checks (90-minute timeout, specific env vars like NEMOCLAW_AGENT=openclaw, COMPATIBLE_API_KEY staging, fake token requirements, test file path, artifact upload configuration) are not enforced.
  • Required action: Add validateOpenClawChannelsAddRemoveJob and validateOpenClawChannelsStopStartJob functions mirroring validateOpenClawChannelsCredentialRewriteJob pattern. Call them from validateE2eWorkflowBoundary.
  • Expected follow-up: Fix before merge or get explicit maintainer override.
  • Verification: grep for 'validateOpenClawChannelsAddRemoveJob' and 'validateOpenClawChannelsStopStartJob' in tools/e2e/workflow-boundary.mts; verify they are called in validateE2eWorkflowBoundary.
  • Missing regression test: test/e2e/support/openclaw-channels-add-remove-workflow-boundary.test.ts and openclaw-channels-stop-start-workflow-boundary.test.ts
  • Done when: The required change is committed and verification passes: grep for 'validateOpenClawChannelsAddRemoveJob' and 'validateOpenClawChannelsStopStartJob' in tools/e2e/workflow-boundary.mts; verify they are called in validateE2eWorkflowBoundary.
  • Evidence: Only validateOpenClawChannelsCredentialRewriteJob, validateOpenClawChannelsConflictGuardJob, validateOpenClawChannelsPairingJob, validateOpenClawChannelsTelegramInjectionSafetyJob exist
Review findings by urgency: 8 required fixes, 13 items to resolve/justify, 3 in-scope improvements

⚠️ Resolve or justify before merge

Investigate these in the current review; either fix them, explain why they are not applicable, or document the accepted risk.

PRA-1 Resolve/justify — Source-of-truth review needed: test/e2e/live/channels-lifecycle-helpers.ts:359

  • Location: not file-specific
  • Category: architecture
  • Problem: The advisor marked localized patch analysis as missing.
  • Impact: A localized workaround can preserve or hide an invalid state when the source boundary is unclear.
  • Recommended action: Identify the invalid state, source boundary, source-fix constraint, regression test, and removal condition before merging the localized behavior.
  • Expected follow-up: Resolve in this PR or explain why the risk is acceptable.
  • Verification: Inspect the localized patch and source-of-truth review fields for a concrete invalid state, source boundary, source-fix constraint, regression test, and removal condition.
  • Missing regression test: expectHermesProtocolCredentialRewrite called from runChannelsStopStartTarget for Hermes agent after 'start' phase (line ~1049)
  • Done when: The risk is fixed or explicitly justified in the PR. Verification: Inspect the localized patch and source-of-truth review fields for a concrete invalid state, source boundary, source-fix constraint, regression test, and removal condition.
  • Evidence: hermesChannelProbe at line 359 only greps placeholder patterns; expectHermesProtocolCredentialRewrite at line 618 does real Discord/Slack L7 proxy rewrite proofs but only called in runChannelsStopStartTarget

PRA-2 Resolve/justify — Source-of-truth review needed: src/lib/agent/onboard.ts:1

  • Location: not file-specific
  • Category: architecture
  • Problem: The advisor marked localized patch analysis as missing.
  • Impact: A localized workaround can preserve or hide an invalid state when the source boundary is unclear.
  • Recommended action: Identify the invalid state, source boundary, source-fix constraint, regression test, and removal condition before merging the localized behavior.
  • Expected follow-up: Resolve in this PR or explain why the risk is acceptable.
  • Verification: Inspect the localized patch and source-of-truth review fields for a concrete invalid state, source boundary, source-fix constraint, regression test, and removal condition.
  • Missing regression test: Existing onboard.ts unit/integration tests cover all three areas; extraction should not change behavior
  • Done when: The risk is fixed or explicitly justified in the PR. Verification: Inspect the localized patch and source-of-truth review fields for a concrete invalid state, source boundary, source-fix constraint, regression test, and removal condition.
  • Evidence: ensureAgentBaseImage (64-136), createAgentSandbox (142-177), dashboard functions (467-637) are all self-contained and exported

PRA-3 Resolve/justify — Source-of-truth review needed: src/lib/build-context.ts:27

  • Location: not file-specific
  • Category: architecture
  • Problem: The advisor marked localized patch analysis as missing.
  • Impact: A localized workaround can preserve or hide an invalid state when the source boundary is unclear.
  • Recommended action: Identify the invalid state, source boundary, source-fix constraint, regression test, and removal condition before merging the localized behavior.
  • Expected follow-up: Resolve in this PR or explain why the risk is acceptable.
  • Verification: Inspect the localized patch and source-of-truth review fields for a concrete invalid state, source boundary, source-fix constraint, regression test, and removal condition.
  • Missing regression test: Script validating all four locations match the central source
  • Done when: The risk is fixed or explicitly justified in the PR. Verification: Inspect the localized patch and source-of-truth review fields for a concrete invalid state, source boundary, source-fix constraint, regression test, and removal condition.
  • Evidence: build-context.ts misses .claude, .DS_Store, .env*, *.key; onboard.ts misses .ruff_cache, .pytest_cache, .mypy_cache; .dockerignore adds secrets patterns; .gitignore adds .version, *.tsbuildinfo, .idea/

PRA-4 Resolve/justify — Source-of-truth review needed: tools/e2e/workflow-boundary.mts:1

  • Location: not file-specific
  • Category: architecture
  • Problem: The advisor marked localized patch analysis as missing.
  • Impact: A localized workaround can preserve or hide an invalid state when the source boundary is unclear.
  • Recommended action: Identify the invalid state, source boundary, source-fix constraint, regression test, and removal condition before merging the localized behavior.
  • Expected follow-up: Resolve in this PR or explain why the risk is acceptable.
  • Verification: Inspect the localized patch and source-of-truth review fields for a concrete invalid state, source boundary, source-fix constraint, regression test, and removal condition.
  • Missing regression test: N/A - documentation
  • Done when: The risk is fixed or explicitly justified in the PR. Verification: Inspect the localized patch and source-of-truth review fields for a concrete invalid state, source boundary, source-fix constraint, regression test, and removal condition.
  • Evidence: First 20 lines show only imports and type definitions, no source-of-truth documentation

PRA-5 Resolve/justify — Source-of-truth review needed: test/e2e/live/issue-2478-crash-loop-recovery.test.ts:1

  • Location: not file-specific
  • Category: architecture
  • Problem: The advisor marked localized patch analysis as needs_followup.
  • Impact: A localized workaround can preserve or hide an invalid state when the source boundary is unclear.
  • Recommended action: Identify the invalid state, source boundary, source-fix constraint, regression test, and removal condition before merging the localized behavior.
  • Expected follow-up: Resolve in this PR or explain why the risk is acceptable.
  • Verification: Inspect the localized patch and source-of-truth review fields for a concrete invalid state, source boundary, source-fix constraint, regression test, and removal condition.
  • Missing regression test: Test itself is the regression test for issue [DGX Spark] Gateway crash loop on startup: @homebridge/ciao networkInterfaces() returns EPERM in OpenShell sandbox #2478 crash-loop fix
  • Done when: The risk is fixed or explicitly justified in the PR. Verification: Inspect the localized patch and source-of-truth review fields for a concrete invalid state, source boundary, source-fix constraint, regression test, and removal condition.
  • Evidence: Test onboards sandbox, kills and recovers gateway via production connect --probe-only path, verifies guard-chain preloads remain present, proves inference.local keeps serving models

PRA-13 Resolve/justify — workflow-boundary.mts lacks source-of-truth documentation

  • Location: tools/e2e/workflow-boundary.mts:1
  • Category: architecture
  • Problem: No header comment documents it as the authoritative validator for workflow security properties or references the self-test file (PRA-2).
  • Impact: Maintainers may not recognize this file as the authoritative validator for workflow security properties. Lack of self-test reference makes it harder to verify validator correctness.
  • Recommended action: Add header comment to workflow-boundary.mts documenting it as the source of truth for workflow security properties. Reference the self-test file from PRA-2 once created.
  • Expected follow-up: Resolve in this PR or explain why the risk is acceptable.
  • Verification: Check if tools/e2e/workflow-boundary.mts has a header comment documenting its role as source of truth for workflow security properties.
  • Missing regression test: None (documentation)
  • Done when: The risk is fixed or explicitly justified in the PR. Verification: Check if tools/e2e/workflow-boundary.mts has a header comment documenting its role as source of truth for workflow security properties.
  • Evidence: First 20 lines of tools/e2e/workflow-boundary.mts show imports but no source-of-truth documentation header

PRA-14 Resolve/justify — channels-lifecycle-helpers.ts is a 1059-line monolith without architecture justification

  • Location: test/e2e/live/channels-lifecycle-helpers.ts:1
  • Category: architecture
  • Problem: Contains registry helpers, provider helpers, policy helpers, config helpers, protocol helpers all mixed together. No header comment explains why the monolith is correct.
  • Impact: High coupling makes changes risky. Hard to understand, test, and maintain. Cross-lifecycle refactoring requires touching this single large file.
  • Recommended action: Either split into channels-registry-helpers.ts, channels-provider-helpers.ts, channels-policy-helpers.ts, channels-config-helpers.ts, channels-protocol-helpers.ts if shared assertions are stable, OR add a header comment explaining why the monolith is correct (e.g., frequent cross-lifecycle refactoring, tightly coupled state machine).
  • Expected follow-up: Resolve in this PR or explain why the risk is acceptable.
  • Verification: Check if channels-lifecycle-helpers.ts has a header comment justifying the monolith, or if it has been split into smaller focused helpers.
  • Missing regression test: None (architecture)
  • Done when: The risk is fixed or explicitly justified in the PR. Verification: Check if channels-lifecycle-helpers.ts has a header comment justifying the monolith, or if it has been split into smaller focused helpers.
  • Evidence: File is 1059 lines with no architecture justification comment at top

PRA-15 Resolve/justify — hermesChannelProbe workaround lacks documentation and removal condition

  • Location: test/e2e/live/channels-lifecycle-helpers.ts:359
  • Category: architecture
  • Problem: hermesChannelProbe probes /sandbox/.hermes/.env with grep for placeholder patterns instead of using the proper credential rewrite verification (expectHermesProtocolCredentialRewrite). Used in agentConfigContains for Hermes agent.
  • Impact: Workaround may become permanent technical debt. Without explicit removal condition, future maintainers won't know when it's safe to delete. The real verification (expectHermesProtocolCredentialRewrite at line ~1049) is only called in runChannelsStopStartTarget for Hermes after 'start' phase.
  • Recommended action: Document hermesChannelProbe as a known workaround with explicit removal condition in code comments. Reference expectHermesProtocolCredentialRewrite as the regression guard that will replace it. Call expectHermesProtocolCredentialRewrite from runChannelsAddRemoveTarget for Hermes agent as well.
  • Expected follow-up: Resolve in this PR or explain why the risk is acceptable.
  • Verification: Check if hermesChannelProbe function has a comment documenting it as a workaround with explicit removal condition referencing expectHermesProtocolCredentialRewrite.
  • Missing regression test: None (documentation of existing workaround)
  • Done when: The risk is fixed or explicitly justified in the PR. Verification: Check if hermesChannelProbe function has a comment documenting it as a workaround with explicit removal condition referencing expectHermesProtocolCredentialRewrite.
  • Evidence: hermesChannelProbe function at line 359 has no workaround documentation; expectHermesProtocolCredentialRewrite called at line 1049 in runChannelsStopStartTarget only

PRA-16 Resolve/justify — src/lib/agent/onboard.ts is a 637-line monolith with no helper extraction

  • Location: src/lib/agent/onboard.ts:1
  • Category: architecture
  • Problem: Contains ensureAgentBaseImage (lines 64-136), createAgentSandbox (lines 142-177), getAgentDashboardInfo/printDashboardUi/printAdditionalForwardPorts/resolveUrlPort (lines 467-637). No helper extraction.
  • Impact: High coupling, hard to test in isolation, exceeds hotspot threshold. Changes to base image logic, sandbox factory, or dashboard printing all touch the same file.
  • Recommended action: Extract cohesive helpers before merge: agent-base-image.ts (ensureAgentBaseImage), agent-sandbox-factory.ts (createAgentSandbox), agent-dashboard.ts (getAgentDashboardInfo, printDashboardUi, printAdditionalForwardPorts, resolveUrlPort). Update onboard.ts to import from them.
  • Expected follow-up: Resolve in this PR or explain why the risk is acceptable.
  • Verification: Check if src/lib/agent/agent-base-image.ts, agent-sandbox-factory.ts, agent-dashboard.ts exist and onboard.ts imports from them.
  • Missing regression test: None (refactoring)
  • Done when: The risk is fixed or explicitly justified in the PR. Verification: Check if src/lib/agent/agent-base-image.ts, agent-sandbox-factory.ts, agent-dashboard.ts exist and onboard.ts imports from them.
  • Evidence: File is 637 lines with three distinct functional areas not extracted

PRA-17 Resolve/justify — Build context exclusion patterns duplicated across 4 files with imperfect synchronization

  • Location: src/lib/build-context.ts:27
  • Category: architecture
  • Problem: EXCLUDED_SEGMENTS in build-context.ts (7 patterns), createAgentSandbox filter in onboard.ts (5 patterns: .git, .venv, __pycache__, node_modules, .claude), .dockerignore (14+ patterns), .gitignore (15+ patterns) all diverge. No single source of truth.
  • Impact: Divergent exclusion lists can cause: sensitive files leaking into Docker build context (if .dockerignore is stricter), unnecessary files bloating build context (if build-context.ts is looser), or build failures (if patterns differ).
  • Recommended action: Centralize exclusion patterns in a single source file (e.g., src/lib/core/build-context-exclusions.ts) that both build-context.ts and a sync script for .dockerignore/.gitignore can reference. At minimum, add a comment in build-context.ts referencing .dockerignore and .gitignore to keep them in sync.
  • Expected follow-up: Resolve in this PR or explain why the risk is acceptable.
  • Verification: Compare EXCLUDED_SEGMENTS in build-context.ts, createAgentSandbox filter in onboard.ts, .dockerignore, and .gitignore for consistency.
  • Missing regression test: Test or script that validates exclusion pattern consistency across all four locations
  • Done when: The risk is fixed or explicitly justified in the PR. Verification: Compare EXCLUDED_SEGMENTS in build-context.ts, createAgentSandbox filter in onboard.ts, .dockerignore, and .gitignore for consistency.
  • Evidence: build-context.ts EXCLUDED_SEGMENTS = {.venv, .ruff_cache, .pytest_cache, .mypy_cache, __pycache__, node_modules, .git}; onboard.ts filter = [.git, .venv, __pycache__, node_modules, .claude]; .dockerignore adds .DS_Store, .env*, *.key, *.pem, etc.; .gitignore adds .version, *.tsbuildinfo, .idea/, etc.

PRA-18 Resolve/justify — Missing hermes-channels-conflict-guard job in CI workflow

  • Location: .github/workflows/e2e.yaml:3342
  • Category: workflow
  • Problem: Same as PRA-7 but categorized as workflow issue.
  • Impact: Hermes agent duplicate credential conflict detection has no CI coverage. Credential conflict bugs in Hermes path would go undetected.
  • Recommended action: Add hermes-channels-conflict-guard job to e2e.yaml mirroring openclaw-channels-conflict-guard. Add to report-to-pr needs list.
  • Expected follow-up: Resolve in this PR or explain why the risk is acceptable.
  • Verification: grep for 'hermes-channels-conflict-guard' in .github/workflows/e2e.yaml; verify it's in the report-to-pr needs array.
  • Missing regression test: hermes-channels-conflict-guard job in e2e.yaml with report-to-pr dependency
  • Done when: The risk is fixed or explicitly justified in the PR. Verification: grep for 'hermes-channels-conflict-guard' in .github/workflows/e2e.yaml; verify it's in the report-to-pr needs array.
  • Evidence: report-to-pr needs array at line 3342 includes openclaw-channels-conflict-guard but not hermes-channels-conflict-guard

PRA-20 Resolve/justify — channels-lifecycle-workflow-boundary.test.ts only tests generic inventory validation

  • Location: test/e2e/support/channels-lifecycle-workflow-boundary.test.ts:1
  • Category: tests
  • Problem: The test only tests generic inventory validation for openclaw-channels-stop-start. It does not test job-specific validators because they don't exist (PRA-18). The test mutates the workflow and expects validateE2eWorkflowBoundary to catch errors, but without job-specific validators, many security properties are not checked.
  • Impact: False confidence: the test passes but doesn't actually verify job-specific security properties for channels add/remove or stop/start jobs.
  • Recommended action: Once PRA-18 and PRA-3 are resolved (job-specific validators added), update channels-lifecycle-workflow-boundary.test.ts to test the new validators, or add dedicated test files for each job.
  • Expected follow-up: Resolve in this PR or explain why the risk is acceptable.
  • Verification: Check if channels-lifecycle-workflow-boundary.test.ts tests job-specific validators for all four channels lifecycle jobs (openclaw/hermes × add-remove/stop-start).
  • Missing regression test: Dedicated workflow boundary tests for each of the four channels lifecycle jobs
  • Done when: The risk is fixed or explicitly justified in the PR. Verification: Check if channels-lifecycle-workflow-boundary.test.ts tests job-specific validators for all four channels lifecycle jobs (openclaw/hermes × add-remove/stop-start).
  • Evidence: Test only validates timeout-minutes, NEMOCLAW_SANDBOX_NAME, DOCKER_CONFIG, NVIDIA_INFERENCE_API_KEY, checkout persist-credentials, install-openshell.sh env -u, run step tokens, test file path, upload-artifact action pinning, name, include-hidden-files, retention-days

PRA-21 Resolve/justify — Inconsistent COMPATIBLE_API_KEY and NVIDIA_INFERENCE_API_KEY handling across channels jobs

  • Location: .github/workflows/e2e.yaml:1
  • Category: security
  • Problem: openclaw-channels-add-remove and hermes-channels-add-remove set COMPATIBLE_API_KEY at run step level. openclaw-channels-stop-start and hermes-channels-stop-start don't set COMPATIBLE_API_KEY at all. openclaw-channels-credential-rewrite doesn't set it either.
  • Impact: Inconsistent credential handling across channels jobs. Some jobs may not properly isolate Docker auth or inference credentials.
  • Recommended action: Standardize COMPATIBLE_API_KEY and NVIDIA_INFERENCE_API_KEY handling across all 6 channels jobs. Ensure job-level env doesn't expose secrets and step-level env uses secrets correctly.
  • Expected follow-up: Resolve in this PR or explain why the risk is acceptable.
  • Verification: Compare env configuration across all 6 channels jobs (openclaw/hermes × add-remove/stop-start/credential-rewrite).
  • Missing regression test: Workflow boundary validators that enforce consistent credential handling (PRA-3, PRA-18)
  • Done when: The risk is fixed or explicitly justified in the PR. Verification: Compare env configuration across all 6 channels jobs (openclaw/hermes × add-remove/stop-start/credential-rewrite).
  • Evidence: openclaw-channels-add-remove sets COMPATIBLE_API_KEY at run step (line 3841); openclaw-channels-stop-start has no COMPATIBLE_API_KEY (line 4075); openclaw-channels-credential-rewrite has no COMPATIBLE_API_KEY (line 3031)

💡 In-scope improvements

These are lower-risk, not throwaway. Prefer fixing them in this PR when they are local to changed code; defer only with rationale or a linked follow-up.

PRA-22 Improvement — ensureAgentBaseImage function can be extracted to agent-base-image.ts helper

  • Location: src/lib/agent/onboard.ts:150
  • Category: architecture
  • Problem: ensureAgentBaseImage function (lines 64-136) can be extracted to agent-base-image.ts helper.
  • Impact: Reduces onboard.ts size by ~73 lines. Enables isolated testing of base image resolution logic.
  • Suggested action: Create src/lib/agent/agent-base-image.ts with ensureAgentBaseImage and resolveSandboxBaseImage exports. Update onboard.ts to import from it.
  • Expected follow-up: Prefer a current-PR fix when local to changed code; defer only with rationale or linked follow-up.
  • Verification: Check if src/lib/agent/agent-base-image.ts exists and is imported by onboard.ts.
  • Missing regression test: None (refactoring)
  • Done when: The local improvement is applied, or the PR notes why it should be deferred.
  • Evidence: Function is self-contained at lines 64-136 with no dependencies on other onboard.ts internals

PRA-23 Improvement — createAgentSandbox function can be extracted to agent-sandbox-factory.ts helper

  • Location: src/lib/agent/onboard.ts:160
  • Category: architecture
  • Problem: createAgentSandbox function (lines 142-177) can be extracted to agent-sandbox-factory.ts helper.
  • Impact: Reduces onboard.ts size by ~36 lines. Enables isolated testing of build context staging logic.
  • Suggested action: Create src/lib/agent/agent-sandbox-factory.ts with createAgentSandbox export. Update onboard.ts to import from it.
  • Expected follow-up: Prefer a current-PR fix when local to changed code; defer only with rationale or linked follow-up.
  • Verification: Check if src/lib/agent/agent-sandbox-factory.ts exists and is imported by onboard.ts.
  • Missing regression test: None (refactoring)
  • Done when: The local improvement is applied, or the PR notes why it should be deferred.
  • Evidence: Function is self-contained at lines 142-177, uses ensureAgentBaseImage which would be imported from agent-base-image.ts

PRA-24 Improvement — Dashboard printing functions can be extracted to agent-dashboard.ts helper

  • Location: src/lib/agent/onboard.ts:550
  • Category: architecture
  • Problem: Dashboard printing functions (getAgentDashboardInfo, printDashboardUi, printAdditionalForwardPorts, resolveUrlPort at lines 467-637) can be extracted to agent-dashboard.ts helper.
  • Impact: Reduces onboard.ts size by ~170 lines. Enables isolated testing of dashboard URL generation and printing logic.
  • Suggested action: Create src/lib/agent/agent-dashboard.ts with getAgentDashboardInfo, printDashboardUi, printAdditionalForwardPorts, resolveUrlPort exports. Update onboard.ts to import from it.
  • Expected follow-up: Prefer a current-PR fix when local to changed code; defer only with rationale or linked follow-up.
  • Verification: Check if src/lib/agent/agent-dashboard.ts exists and is imported by onboard.ts.
  • Missing regression test: None (refactoring)
  • Done when: The local improvement is applied, or the PR notes why it should be deferred.
  • Evidence: Four functions at lines 467-637 form cohesive dashboard printing module with no external dependencies
Simplification opportunities: 3 possible cuts, net -279 lines possible

These are safe simplification checks only. Do not remove validation, security controls, data-loss prevention, or required tests.

  • PRA-22 shrink (src/lib/agent/onboard.ts:150): ensureAgentBaseImage function (lines 64-136) from src/lib/agent/onboard.ts
    • Replacement: import { ensureAgentBaseImage } from './agent-base-image.ts'
    • Net: -73 lines
    • Safety boundary: Must preserve base image resolution logic for both OpenClaw and Hermes agents; function is already exported and used by createAgentSandbox
  • PRA-23 shrink (src/lib/agent/onboard.ts:160): createAgentSandbox function (lines 142-177) from src/lib/agent/onboard.ts
    • Replacement: import { createAgentSandbox } from './agent-sandbox-factory.ts'
    • Net: -36 lines
    • Safety boundary: Must preserve Dockerfile staging logic and base image ARG replacement; function is already exported
  • PRA-24 shrink (src/lib/agent/onboard.ts:550): getAgentDashboardInfo, printDashboardUi, printAdditionalForwardPorts, resolveUrlPort functions (lines 467-637) from src/lib/agent/onboard.ts
    • Replacement: import { getAgentDashboardInfo, printDashboardUi, printAdditionalForwardPorts, resolveUrlPort } from './agent-dashboard.ts'
    • Net: -170 lines
    • Safety boundary: Must preserve dashboard URL generation, token redaction, and forward port printing logic; all functions already exported
Test follow-ups to resolve or justify

If these cover changed behavior, prefer adding them in this PR; otherwise state why existing coverage is enough or link the follow-up.

  • PRA-T1 Runtime validation — workflow-boundary-validator self-test: validateE2eWorkflowBoundary with valid/invalid YAML fixtures. Runtime/sandbox/infrastructure paths need behavioral runtime validation: .github/workflows/brev-nightly-e2e.yaml, .github/workflows/e2e-branch-validation.yaml, .github/workflows/e2e.yaml, .github/workflows/regression-e2e.yaml, docs/manage-sandboxes/messaging-channels.mdx, docs/reference/troubleshooting.mdx, src/lib/actions/sandbox/policy-explain.ts, tools/e2e/upload-e2e-artifacts-workflow-boundary.mts.
  • PRA-T2 Runtime validation — workflow-boundary-validator self-test: evaluateE2eWorkflowDispatchSelectors selector parsing and unknown job rejection. Runtime/sandbox/infrastructure paths need behavioral runtime validation: .github/workflows/brev-nightly-e2e.yaml, .github/workflows/e2e-branch-validation.yaml, .github/workflows/e2e.yaml, .github/workflows/regression-e2e.yaml, docs/manage-sandboxes/messaging-channels.mdx, docs/reference/troubleshooting.mdx, src/lib/actions/sandbox/policy-explain.ts, tools/e2e/upload-e2e-artifacts-workflow-boundary.mts.
  • PRA-T3 Runtime validation — workflow-boundary-validator self-test: validateFreeStandingWorkflowInventory duplicate detection and missing keys. Runtime/sandbox/infrastructure paths need behavioral runtime validation: .github/workflows/brev-nightly-e2e.yaml, .github/workflows/e2e-branch-validation.yaml, .github/workflows/e2e.yaml, .github/workflows/regression-e2e.yaml, docs/manage-sandboxes/messaging-channels.mdx, docs/reference/troubleshooting.mdx, src/lib/actions/sandbox/policy-explain.ts, tools/e2e/upload-e2e-artifacts-workflow-boundary.mts.
  • PRA-T4 Runtime validation — workflow-boundary-validator self-test: readFreeStandingJobsInventory caching and error propagation. Runtime/sandbox/infrastructure paths need behavioral runtime validation: .github/workflows/brev-nightly-e2e.yaml, .github/workflows/e2e-branch-validation.yaml, .github/workflows/e2e.yaml, .github/workflows/regression-e2e.yaml, docs/manage-sandboxes/messaging-channels.mdx, docs/reference/troubleshooting.mdx, src/lib/actions/sandbox/policy-explain.ts, tools/e2e/upload-e2e-artifacts-workflow-boundary.mts.
  • PRA-T5 Runtime validation — validateHermesChannelsAddRemoveJob security properties. Runtime/sandbox/infrastructure paths need behavioral runtime validation: .github/workflows/brev-nightly-e2e.yaml, .github/workflows/e2e-branch-validation.yaml, .github/workflows/e2e.yaml, .github/workflows/regression-e2e.yaml, docs/manage-sandboxes/messaging-channels.mdx, docs/reference/troubleshooting.mdx, src/lib/actions/sandbox/policy-explain.ts, tools/e2e/upload-e2e-artifacts-workflow-boundary.mts.
  • PRA-T6 workflow-boundary.mts trusted validator lacks self-test coverage — Add test/e2e/support/workflow-boundary-validator.test.ts with fixture-based tests covering: validateE2eWorkflowBoundary (valid/invalid workflow YAML), evaluateE2eWorkflowDispatchSelectors (selector parsing, unknown job rejection), validateFreeStandingWorkflowInventory (duplicate detection, missing keys), and readFreeStandingJobsInventory (caching, error propagation).
  • PRA-T7 No Hermes workflow boundary validator tests exist — Add test/e2e/support/hermes-channels-add-remove-workflow-boundary.test.ts and test/e2e/support/hermes-channels-stop-start-workflow-boundary.test.ts mirroring the OpenClaw pattern.
  • PRA-T8 Missing hermes-channels-conflict-guard workflow job — Add hermes-channels-conflict-guard job to e2e.yaml mirroring openclaw-channels-conflict-guard. Add to report-to-pr needs list.
Since last review details

Current findings, using the urgency labels above:

PRA-1 Resolve/justify — Source-of-truth review needed: test/e2e/live/channels-lifecycle-helpers.ts:359

  • Location: not file-specific
  • Category: architecture
  • Problem: The advisor marked localized patch analysis as missing.
  • Impact: A localized workaround can preserve or hide an invalid state when the source boundary is unclear.
  • Recommended action: Identify the invalid state, source boundary, source-fix constraint, regression test, and removal condition before merging the localized behavior.
  • Expected follow-up: Resolve in this PR or explain why the risk is acceptable.
  • Verification: Inspect the localized patch and source-of-truth review fields for a concrete invalid state, source boundary, source-fix constraint, regression test, and removal condition.
  • Missing regression test: expectHermesProtocolCredentialRewrite called from runChannelsStopStartTarget for Hermes agent after 'start' phase (line ~1049)
  • Done when: The risk is fixed or explicitly justified in the PR. Verification: Inspect the localized patch and source-of-truth review fields for a concrete invalid state, source boundary, source-fix constraint, regression test, and removal condition.
  • Evidence: hermesChannelProbe at line 359 only greps placeholder patterns; expectHermesProtocolCredentialRewrite at line 618 does real Discord/Slack L7 proxy rewrite proofs but only called in runChannelsStopStartTarget

PRA-2 Resolve/justify — Source-of-truth review needed: src/lib/agent/onboard.ts:1

  • Location: not file-specific
  • Category: architecture
  • Problem: The advisor marked localized patch analysis as missing.
  • Impact: A localized workaround can preserve or hide an invalid state when the source boundary is unclear.
  • Recommended action: Identify the invalid state, source boundary, source-fix constraint, regression test, and removal condition before merging the localized behavior.
  • Expected follow-up: Resolve in this PR or explain why the risk is acceptable.
  • Verification: Inspect the localized patch and source-of-truth review fields for a concrete invalid state, source boundary, source-fix constraint, regression test, and removal condition.
  • Missing regression test: Existing onboard.ts unit/integration tests cover all three areas; extraction should not change behavior
  • Done when: The risk is fixed or explicitly justified in the PR. Verification: Inspect the localized patch and source-of-truth review fields for a concrete invalid state, source boundary, source-fix constraint, regression test, and removal condition.
  • Evidence: ensureAgentBaseImage (64-136), createAgentSandbox (142-177), dashboard functions (467-637) are all self-contained and exported

PRA-3 Resolve/justify — Source-of-truth review needed: src/lib/build-context.ts:27

  • Location: not file-specific
  • Category: architecture
  • Problem: The advisor marked localized patch analysis as missing.
  • Impact: A localized workaround can preserve or hide an invalid state when the source boundary is unclear.
  • Recommended action: Identify the invalid state, source boundary, source-fix constraint, regression test, and removal condition before merging the localized behavior.
  • Expected follow-up: Resolve in this PR or explain why the risk is acceptable.
  • Verification: Inspect the localized patch and source-of-truth review fields for a concrete invalid state, source boundary, source-fix constraint, regression test, and removal condition.
  • Missing regression test: Script validating all four locations match the central source
  • Done when: The risk is fixed or explicitly justified in the PR. Verification: Inspect the localized patch and source-of-truth review fields for a concrete invalid state, source boundary, source-fix constraint, regression test, and removal condition.
  • Evidence: build-context.ts misses .claude, .DS_Store, .env*, *.key; onboard.ts misses .ruff_cache, .pytest_cache, .mypy_cache; .dockerignore adds secrets patterns; .gitignore adds .version, *.tsbuildinfo, .idea/

PRA-4 Resolve/justify — Source-of-truth review needed: tools/e2e/workflow-boundary.mts:1

  • Location: not file-specific
  • Category: architecture
  • Problem: The advisor marked localized patch analysis as missing.
  • Impact: A localized workaround can preserve or hide an invalid state when the source boundary is unclear.
  • Recommended action: Identify the invalid state, source boundary, source-fix constraint, regression test, and removal condition before merging the localized behavior.
  • Expected follow-up: Resolve in this PR or explain why the risk is acceptable.
  • Verification: Inspect the localized patch and source-of-truth review fields for a concrete invalid state, source boundary, source-fix constraint, regression test, and removal condition.
  • Missing regression test: N/A - documentation
  • Done when: The risk is fixed or explicitly justified in the PR. Verification: Inspect the localized patch and source-of-truth review fields for a concrete invalid state, source boundary, source-fix constraint, regression test, and removal condition.
  • Evidence: First 20 lines show only imports and type definitions, no source-of-truth documentation

PRA-5 Resolve/justify — Source-of-truth review needed: test/e2e/live/issue-2478-crash-loop-recovery.test.ts:1

  • Location: not file-specific
  • Category: architecture
  • Problem: The advisor marked localized patch analysis as needs_followup.
  • Impact: A localized workaround can preserve or hide an invalid state when the source boundary is unclear.
  • Recommended action: Identify the invalid state, source boundary, source-fix constraint, regression test, and removal condition before merging the localized behavior.
  • Expected follow-up: Resolve in this PR or explain why the risk is acceptable.
  • Verification: Inspect the localized patch and source-of-truth review fields for a concrete invalid state, source boundary, source-fix constraint, regression test, and removal condition.
  • Missing regression test: Test itself is the regression test for issue [DGX Spark] Gateway crash loop on startup: @homebridge/ciao networkInterfaces() returns EPERM in OpenShell sandbox #2478 crash-loop fix
  • Done when: The risk is fixed or explicitly justified in the PR. Verification: Inspect the localized patch and source-of-truth review fields for a concrete invalid state, source boundary, source-fix constraint, regression test, and removal condition.
  • Evidence: Test onboards sandbox, kills and recovers gateway via production connect --probe-only path, verifies guard-chain preloads remain present, proves inference.local keeps serving models

PRA-6 Required — workflow-boundary.mts trusted validator lacks self-test coverage

  • Location: tools/e2e/workflow-boundary.mts:1
  • Category: tests
  • Problem: The workflow-boundary.mts validator enforces security properties (action SHA pinning, secret non-exposure, checkout persist-credentials=false, artifact path hygiene, DOCKER_CONFIG handling) for all free-standing E2E jobs but has no tests validating its own logic. A bug in the validator could silently allow unsafe workflow patterns.
  • Impact: Trusted code boundary without regression tests; validator drift could permit unpinned actions, secret leakage, or unsafe artifact paths in production workflows.
  • Required action: Add test/e2e/support/workflow-boundary-validator.test.ts with fixture-based tests covering: validateE2eWorkflowBoundary (valid/invalid workflow YAML), evaluateE2eWorkflowDispatchSelectors (selector parsing, unknown job rejection), validateFreeStandingWorkflowInventory (duplicate detection, missing keys), and readFreeStandingJobsInventory (caching, error propagation).
  • Expected follow-up: Fix before merge or get explicit maintainer override.
  • Verification: Check if test/e2e/support/workflow-boundary-validator.test.ts exists and covers the four validator functions with valid/invalid fixtures.
  • Missing regression test: test/e2e/support/workflow-boundary-validator.test.ts with fixture-based tests for all four exported validator functions
  • Done when: The required change is committed and verification passes: Check if test/e2e/support/workflow-boundary-validator.test.ts exists and covers the four validator functions with valid/invalid fixtures.
  • Evidence: tools/e2e/workflow-boundary.mts exports validateE2eWorkflowBoundary, evaluateE2eWorkflowDispatchSelectors, validateFreeStandingWorkflowInventory, readFreeStandingJobsInventory but no test file exercises them

PRA-7 Required — Missing validateHermesChannelsAddRemove and validateHermesChannelsStopStart validators

  • Location: tools/e2e/workflow-boundary.mts:1
  • Category: security
  • Problem: The hermes-channels-add-remove and hermes-channels-stop-start jobs exist in e2e.yaml but have no dedicated validator functions checking: hosted inference config, fake tokens, Docker auth isolation (RUNNER_TEMP), test file path, artifact upload path, checkout persist-credentials=false, action SHA pinning.
  • Impact: Hermes channels jobs bypass job-specific security checks. Only generic inventory validation applies, leaving fake token enforcement, Docker auth isolation, checkout safety, and action pinning unchecked for these jobs.
  • Required action: Add validateHermesChannelsAddRemove and validateHermesChannelsStopStart functions mirroring the OpenClaw validators, checking all security properties. Call them from validateE2eWorkflowBoundary.
  • Expected follow-up: Fix before merge or get explicit maintainer override.
  • Verification: grep for 'validateHermesChannelsAddRemove' and 'validateHermesChannelsStopStart' in tools/e2e/workflow-boundary.mts; verify they are called in validateE2eWorkflowBoundary.
  • Missing regression test: test/e2e/support/hermes-channels-add-remove-workflow-boundary.test.ts and hermes-channels-stop-start-workflow-boundary.test.ts mirroring openclaw-channels-pairing-workflow-boundary.test.ts pattern
  • Done when: The required change is committed and verification passes: grep for 'validateHermesChannelsAddRemove' and 'validateHermesChannelsStopStart' in tools/e2e/workflow-boundary.mts; verify they are called in validateE2eWorkflowBoundary.
  • Evidence: tools/e2e/workflow-boundary.mts only has validators for openclaw-channels-credential-rewrite, openclaw-channels-conflict-guard, openclaw-channels-pairing, openclaw-channels-telegram-injection-safety — no Hermes channels add-remove/stop-start validators

PRA-8 Required — No Hermes workflow boundary validator tests exist

  • Location: test/e2e/support:1
  • Category: tests
  • Problem: The OpenClaw pattern has test/e2e/support/openclaw-channels-pairing-workflow-boundary.test.ts that mutates the workflow and asserts validateE2eWorkflowBoundary catches security violations. No equivalent exists for Hermes channels jobs.
  • Impact: Workflow boundary drift for Hermes channels jobs would not be caught by tests. Security properties like fake token enforcement, Docker auth isolation, and action pinning have no regression test coverage.
  • Required action: Add test/e2e/support/hermes-channels-add-remove-workflow-boundary.test.ts and test/e2e/support/hermes-channels-stop-start-workflow-boundary.test.ts mirroring the OpenClaw pattern.
  • Expected follow-up: Fix before merge or get explicit maintainer override.
  • Verification: Check if test/e2e/support/hermes-channels-add-remove-workflow-boundary.test.ts and hermes-channels-stop-start-workflow-boundary.test.ts exist.
  • Missing regression test: test/e2e/support/hermes-channels-add-remove-workflow-boundary.test.ts and hermes-channels-stop-start-workflow-boundary.test.ts
  • Done when: The required change is committed and verification passes: Check if test/e2e/support/hermes-channels-add-remove-workflow-boundary.test.ts and hermes-channels-stop-start-workflow-boundary.test.ts exist.
  • Evidence: Only test/e2e/support/hermes-workflow-boundary.test.ts exists, which only tests hermes-e2e job model pinning, not channels add-remove/stop-start jobs

PRA-9 Required — Missing Hermes equivalent of credential rewrite test

  • Location: test/e2e/live:1
  • Category: security
  • Problem: OpenClaw has test/e2e/live/openclaw-channels-credential-rewrite.test.ts covering: provider placeholder verification (/sandbox/.hermes/.env), L7 proxy rewrite proofs for Discord Gateway (WebSocket), Slack REST (auth.test, apps.connections.open), Slack Socket Mode (HELLO/EVENT/ACK), and raw token leak checks for Hermes config paths. No Hermes version exists.
  • Impact: Hermes agent credential rewrite path (L7 proxy token rewriting for Discord/Slack) has no E2E coverage. Token leakage or rewrite failures in Hermes config paths (/sandbox/.hermes/.env) would go undetected.
  • Required action: Add hermes-channels-credential-rewrite.test.ts covering Hermes provider placeholder verification (/sandbox/.hermes/.env), L7 proxy rewrite proofs for Discord Gateway (WebSocket), Slack REST (auth.test, apps.connections.open), Slack Socket Mode (HELLO/EVENT/ACK), and raw token leak checks for Hermes config paths.
  • Expected follow-up: Fix before merge or get explicit maintainer override.
  • Verification: Check if test/e2e/live/hermes-channels-credential-rewrite.test.ts exists with Discord Gateway WebSocket, Slack REST, and Slack Socket Mode rewrite proofs for Hermes.
  • Missing regression test: test/e2e/live/hermes-channels-credential-rewrite.test.ts with Discord Gateway, Slack REST, Slack Socket Mode, and raw token leak proofs for Hermes
  • Done when: The required change is committed and verification passes: Check if test/e2e/live/hermes-channels-credential-rewrite.test.ts exists with Discord Gateway WebSocket, Slack REST, and Slack Socket Mode rewrite proofs for Hermes.
  • Evidence: No test/e2e/live/hermes-channels-credential-rewrite.test.ts file exists; only openclaw-channels-credential-rewrite.test.ts

PRA-10 Required — Missing Hermes pairing tests for Discord Gateway and Slack Socket Mode

  • Location: test/e2e/live:1
  • Category: correctness
  • Problem: OpenClaw has openclaw-channels-pairing-discord.ts and openclaw-channels-pairing-slack.ts using startFakeDiscordGateway/startFakeSlackApi, applyFakePolicy with websocket rewrite, runDiscordGatewayProof/runSlackSocketModeProof with capture assertions, issuePairingRequest, approveAndAssertPairing. No Hermes equivalents exist.
  • Impact: Hermes agent Discord/Slack pairing flows (connect-shell approval, WebSocket/Socket Mode credential rewrite) have no E2E coverage. Pairing failures or token leakage in Hermes paths would go undetected.
  • Required action: Add hermes-channels-pairing-discord.ts and hermes-channels-pairing-slack.ts mirroring OpenClaw patterns. Add corresponding workflow jobs in e2e.yaml.
  • Expected follow-up: Fix before merge or get explicit maintainer override.
  • Verification: Check if test/e2e/live/hermes-channels-pairing-discord.ts and hermes-channels-pairing-slack.ts exist with fake gateway/API servers and rewrite proofs. Verify corresponding jobs in e2e.yaml.
  • Missing regression test: test/e2e/live/hermes-channels-pairing-discord.ts and hermes-channels-pairing-slack.ts with corresponding e2e.yaml jobs
  • Done when: The required change is committed and verification passes: Check if test/e2e/live/hermes-channels-pairing-discord.ts and hermes-channels-pairing-slack.ts exist with fake gateway/API servers and rewrite proofs. Verify corresponding jobs in e2e.yaml.
  • Evidence: No test/e2e/live/hermes-channels-pairing-discord.ts or hermes-channels-pairing-slack.ts files exist; only openclaw-channels-pairing-*

PRA-11 Required — Missing hermes-channels-conflict-guard workflow job

  • Location: .github/workflows/e2e.yaml:3978
  • Category: tests
  • Problem: OpenClaw has openclaw-channels-conflict-guard job in e2e.yaml but no Hermes equivalent.
  • Impact: Hermes agent duplicate credential conflict detection (preventing two sandboxes using the same bot token) has no CI coverage. Credential conflict bugs in Hermes path would go undetected.
  • Required action: Add hermes-channels-conflict-guard job to e2e.yaml mirroring openclaw-channels-conflict-guard. Add to report-to-pr needs list.
  • Expected follow-up: Fix before merge or get explicit maintainer override.
  • Verification: grep for 'hermes-channels-conflict-guard' in .github/workflows/e2e.yaml; verify it's in the report-to-pr needs array.
  • Missing regression test: hermes-channels-conflict-guard job in e2e.yaml with report-to-pr dependency
  • Done when: The required change is committed and verification passes: grep for 'hermes-channels-conflict-guard' in .github/workflows/e2e.yaml; verify it's in the report-to-pr needs array.
  • Evidence: grep shows openclaw-channels-conflict-guard job at line 3339 but no hermes-channels-conflict-guard

PRA-12 Required — Conflict guard test not parameterized over AgentKind

  • Location: test/e2e/live/openclaw-channels-conflict-guard.test.ts:202
  • Category: security
  • Problem: The test only tests 'openclaw' agent (hardcoded in channelPlanWithCredential agent field). Hermes WeChat uses WEIXIN_TOKEN placeholder vs openclaw-weixin for OpenClaw; Hermes Discord/Slack/Telegram/Teams use same keys but different config paths (/sandbox/.hermes/.env vs /sandbox/.openclaw/openclaw.json).
  • Impact: Conflict guard behavior for Hermes agent (WeChat placeholder format, config path differences) is untested. A Hermes-specific conflict guard bug would not be caught.
  • Required action: Parameterize the test loop over both AgentKind and CHANNELS. Update CHANNEL_FIXTURES to include agent-specific placeholder formats: Hermes WeChat uses WEIXIN_TOKEN, OpenClaw uses openclaw-weixin; Hermes Discord/Slack/Telegram/Teams use same keys but different config paths. Loop over ["openclaw", "hermes"] × CREDENTIALED_CHANNELS.
  • Expected follow-up: Fix before merge or get explicit maintainer override.
  • Verification: Check if openclaw-channels-conflict-guard.test.ts loops over AgentKind and uses agent-specific placeholders/config paths.
  • Missing regression test: Parameterized conflict guard test covering both agents × all credentialed channels with agent-specific placeholders
  • Done when: The required change is committed and verification passes: Check if openclaw-channels-conflict-guard.test.ts loops over AgentKind and uses agent-specific placeholders/config paths.
  • Evidence: channelPlanWithCredential hardcodes agent: "openclaw" at line 194; CHANNEL_FIXTURES only defines openclaw placeholders (e.g., wechat placeholder: "openshell:resolve:env:WECHAT_BOT_TOKEN" but Hermes uses WEIXIN_TOKEN)

PRA-13 Resolve/justify — workflow-boundary.mts lacks source-of-truth documentation

  • Location: tools/e2e/workflow-boundary.mts:1
  • Category: architecture
  • Problem: No header comment documents it as the authoritative validator for workflow security properties or references the self-test file (PRA-2).
  • Impact: Maintainers may not recognize this file as the authoritative validator for workflow security properties. Lack of self-test reference makes it harder to verify validator correctness.
  • Recommended action: Add header comment to workflow-boundary.mts documenting it as the source of truth for workflow security properties. Reference the self-test file from PRA-2 once created.
  • Expected follow-up: Resolve in this PR or explain why the risk is acceptable.
  • Verification: Check if tools/e2e/workflow-boundary.mts has a header comment documenting its role as source of truth for workflow security properties.
  • Missing regression test: None (documentation)
  • Done when: The risk is fixed or explicitly justified in the PR. Verification: Check if tools/e2e/workflow-boundary.mts has a header comment documenting its role as source of truth for workflow security properties.
  • Evidence: First 20 lines of tools/e2e/workflow-boundary.mts show imports but no source-of-truth documentation header

PRA-14 Resolve/justify — channels-lifecycle-helpers.ts is a 1059-line monolith without architecture justification

  • Location: test/e2e/live/channels-lifecycle-helpers.ts:1
  • Category: architecture
  • Problem: Contains registry helpers, provider helpers, policy helpers, config helpers, protocol helpers all mixed together. No header comment explains why the monolith is correct.
  • Impact: High coupling makes changes risky. Hard to understand, test, and maintain. Cross-lifecycle refactoring requires touching this single large file.
  • Recommended action: Either split into channels-registry-helpers.ts, channels-provider-helpers.ts, channels-policy-helpers.ts, channels-config-helpers.ts, channels-protocol-helpers.ts if shared assertions are stable, OR add a header comment explaining why the monolith is correct (e.g., frequent cross-lifecycle refactoring, tightly coupled state machine).
  • Expected follow-up: Resolve in this PR or explain why the risk is acceptable.
  • Verification: Check if channels-lifecycle-helpers.ts has a header comment justifying the monolith, or if it has been split into smaller focused helpers.
  • Missing regression test: None (architecture)
  • Done when: The risk is fixed or explicitly justified in the PR. Verification: Check if channels-lifecycle-helpers.ts has a header comment justifying the monolith, or if it has been split into smaller focused helpers.
  • Evidence: File is 1059 lines with no architecture justification comment at top

PRA-15 Resolve/justify — hermesChannelProbe workaround lacks documentation and removal condition

  • Location: test/e2e/live/channels-lifecycle-helpers.ts:359
  • Category: architecture
  • Problem: hermesChannelProbe probes /sandbox/.hermes/.env with grep for placeholder patterns instead of using the proper credential rewrite verification (expectHermesProtocolCredentialRewrite). Used in agentConfigContains for Hermes agent.
  • Impact: Workaround may become permanent technical debt. Without explicit removal condition, future maintainers won't know when it's safe to delete. The real verification (expectHermesProtocolCredentialRewrite at line ~1049) is only called in runChannelsStopStartTarget for Hermes after 'start' phase.
  • Recommended action: Document hermesChannelProbe as a known workaround with explicit removal condition in code comments. Reference expectHermesProtocolCredentialRewrite as the regression guard that will replace it. Call expectHermesProtocolCredentialRewrite from runChannelsAddRemoveTarget for Hermes agent as well.
  • Expected follow-up: Resolve in this PR or explain why the risk is acceptable.
  • Verification: Check if hermesChannelProbe function has a comment documenting it as a workaround with explicit removal condition referencing expectHermesProtocolCredentialRewrite.
  • Missing regression test: None (documentation of existing workaround)
  • Done when: The risk is fixed or explicitly justified in the PR. Verification: Check if hermesChannelProbe function has a comment documenting it as a workaround with explicit removal condition referencing expectHermesProtocolCredentialRewrite.
  • Evidence: hermesChannelProbe function at line 359 has no workaround documentation; expectHermesProtocolCredentialRewrite called at line 1049 in runChannelsStopStartTarget only

PRA-16 Resolve/justify — src/lib/agent/onboard.ts is a 637-line monolith with no helper extraction

  • Location: src/lib/agent/onboard.ts:1
  • Category: architecture
  • Problem: Contains ensureAgentBaseImage (lines 64-136), createAgentSandbox (lines 142-177), getAgentDashboardInfo/printDashboardUi/printAdditionalForwardPorts/resolveUrlPort (lines 467-637). No helper extraction.
  • Impact: High coupling, hard to test in isolation, exceeds hotspot threshold. Changes to base image logic, sandbox factory, or dashboard printing all touch the same file.
  • Recommended action: Extract cohesive helpers before merge: agent-base-image.ts (ensureAgentBaseImage), agent-sandbox-factory.ts (createAgentSandbox), agent-dashboard.ts (getAgentDashboardInfo, printDashboardUi, printAdditionalForwardPorts, resolveUrlPort). Update onboard.ts to import from them.
  • Expected follow-up: Resolve in this PR or explain why the risk is acceptable.
  • Verification: Check if src/lib/agent/agent-base-image.ts, agent-sandbox-factory.ts, agent-dashboard.ts exist and onboard.ts imports from them.
  • Missing regression test: None (refactoring)
  • Done when: The risk is fixed or explicitly justified in the PR. Verification: Check if src/lib/agent/agent-base-image.ts, agent-sandbox-factory.ts, agent-dashboard.ts exist and onboard.ts imports from them.
  • Evidence: File is 637 lines with three distinct functional areas not extracted

PRA-17 Resolve/justify — Build context exclusion patterns duplicated across 4 files with imperfect synchronization

  • Location: src/lib/build-context.ts:27
  • Category: architecture
  • Problem: EXCLUDED_SEGMENTS in build-context.ts (7 patterns), createAgentSandbox filter in onboard.ts (5 patterns: .git, .venv, __pycache__, node_modules, .claude), .dockerignore (14+ patterns), .gitignore (15+ patterns) all diverge. No single source of truth.
  • Impact: Divergent exclusion lists can cause: sensitive files leaking into Docker build context (if .dockerignore is stricter), unnecessary files bloating build context (if build-context.ts is looser), or build failures (if patterns differ).
  • Recommended action: Centralize exclusion patterns in a single source file (e.g., src/lib/core/build-context-exclusions.ts) that both build-context.ts and a sync script for .dockerignore/.gitignore can reference. At minimum, add a comment in build-context.ts referencing .dockerignore and .gitignore to keep them in sync.
  • Expected follow-up: Resolve in this PR or explain why the risk is acceptable.
  • Verification: Compare EXCLUDED_SEGMENTS in build-context.ts, createAgentSandbox filter in onboard.ts, .dockerignore, and .gitignore for consistency.
  • Missing regression test: Test or script that validates exclusion pattern consistency across all four locations
  • Done when: The risk is fixed or explicitly justified in the PR. Verification: Compare EXCLUDED_SEGMENTS in build-context.ts, createAgentSandbox filter in onboard.ts, .dockerignore, and .gitignore for consistency.
  • Evidence: build-context.ts EXCLUDED_SEGMENTS = {.venv, .ruff_cache, .pytest_cache, .mypy_cache, __pycache__, node_modules, .git}; onboard.ts filter = [.git, .venv, __pycache__, node_modules, .claude]; .dockerignore adds .DS_Store, .env*, *.key, *.pem, etc.; .gitignore adds .version, *.tsbuildinfo, .idea/, etc.

PRA-18 Resolve/justify — Missing hermes-channels-conflict-guard job in CI workflow

  • Location: .github/workflows/e2e.yaml:3342
  • Category: workflow
  • Problem: Same as PRA-7 but categorized as workflow issue.
  • Impact: Hermes agent duplicate credential conflict detection has no CI coverage. Credential conflict bugs in Hermes path would go undetected.
  • Recommended action: Add hermes-channels-conflict-guard job to e2e.yaml mirroring openclaw-channels-conflict-guard. Add to report-to-pr needs list.
  • Expected follow-up: Resolve in this PR or explain why the risk is acceptable.
  • Verification: grep for 'hermes-channels-conflict-guard' in .github/workflows/e2e.yaml; verify it's in the report-to-pr needs array.
  • Missing regression test: hermes-channels-conflict-guard job in e2e.yaml with report-to-pr dependency
  • Done when: The risk is fixed or explicitly justified in the PR. Verification: grep for 'hermes-channels-conflict-guard' in .github/workflows/e2e.yaml; verify it's in the report-to-pr needs array.
  • Evidence: report-to-pr needs array at line 3342 includes openclaw-channels-conflict-guard but not hermes-channels-conflict-guard

PRA-19 Required — Missing validators for openclaw-channels-add-remove and openclaw-channels-stop-start jobs

  • Location: tools/e2e/workflow-boundary.mts:1
  • Category: security
  • Problem: These jobs exist in e2e.yaml but have no dedicated validateOpenClawChannelsAddRemoveJob or validateOpenClawChannelsStopStartJob functions. Only generic inventory validation applies.
  • Impact: OpenClaw channels add/remove and stop/start jobs only get generic inventory validation. Job-specific checks (90-minute timeout, specific env vars like NEMOCLAW_AGENT=openclaw, COMPATIBLE_API_KEY staging, fake token requirements, test file path, artifact upload configuration) are not enforced.
  • Required action: Add validateOpenClawChannelsAddRemoveJob and validateOpenClawChannelsStopStartJob functions mirroring validateOpenClawChannelsCredentialRewriteJob pattern. Call them from validateE2eWorkflowBoundary.
  • Expected follow-up: Fix before merge or get explicit maintainer override.
  • Verification: grep for 'validateOpenClawChannelsAddRemoveJob' and 'validateOpenClawChannelsStopStartJob' in tools/e2e/workflow-boundary.mts; verify they are called in validateE2eWorkflowBoundary.
  • Missing regression test: test/e2e/support/openclaw-channels-add-remove-workflow-boundary.test.ts and openclaw-channels-stop-start-workflow-boundary.test.ts
  • Done when: The required change is committed and verification passes: grep for 'validateOpenClawChannelsAddRemoveJob' and 'validateOpenClawChannelsStopStartJob' in tools/e2e/workflow-boundary.mts; verify they are called in validateE2eWorkflowBoundary.
  • Evidence: Only validateOpenClawChannelsCredentialRewriteJob, validateOpenClawChannelsConflictGuardJob, validateOpenClawChannelsPairingJob, validateOpenClawChannelsTelegramInjectionSafetyJob exist

PRA-20 Resolve/justify — channels-lifecycle-workflow-boundary.test.ts only tests generic inventory validation

  • Location: test/e2e/support/channels-lifecycle-workflow-boundary.test.ts:1
  • Category: tests
  • Problem: The test only tests generic inventory validation for openclaw-channels-stop-start. It does not test job-specific validators because they don't exist (PRA-18). The test mutates the workflow and expects validateE2eWorkflowBoundary to catch errors, but without job-specific validators, many security properties are not checked.
  • Impact: False confidence: the test passes but doesn't actually verify job-specific security properties for channels add/remove or stop/start jobs.
  • Recommended action: Once PRA-18 and PRA-3 are resolved (job-specific validators added), update channels-lifecycle-workflow-boundary.test.ts to test the new validators, or add dedicated test files for each job.
  • Expected follow-up: Resolve in this PR or explain why the risk is acceptable.
  • Verification: Check if channels-lifecycle-workflow-boundary.test.ts tests job-specific validators for all four channels lifecycle jobs (openclaw/hermes × add-remove/stop-start).
  • Missing regression test: Dedicated workflow boundary tests for each of the four channels lifecycle jobs
  • Done when: The risk is fixed or explicitly justified in the PR. Verification: Check if channels-lifecycle-workflow-boundary.test.ts tests job-specific validators for all four channels lifecycle jobs (openclaw/hermes × add-remove/stop-start).
  • Evidence: Test only validates timeout-minutes, NEMOCLAW_SANDBOX_NAME, DOCKER_CONFIG, NVIDIA_INFERENCE_API_KEY, checkout persist-credentials, install-openshell.sh env -u, run step tokens, test file path, upload-artifact action pinning, name, include-hidden-files, retention-days

Workflow run details

This is an automated, non-binding review; it still expects maintainers and agents to respond to each required or warning item. Treat suggestions as current-PR improvements when they touch changed code; defer only with maintainer rationale or a linked follow-up. A human maintainer must make the final merge decision.

Signed-off-by: San Dang <sdang@nvidia.com>
@github-actions

Copy link
Copy Markdown
Contributor

Vitest E2E Target Results — ❌ Some jobs failed

Run: 28444218181
Workflow ref: test/e2e-messaging-refactor-6022
Requested targets: (default — all supported)
Requested jobs: (default — all default-enabled free-standing jobs; explicit-only jobs such as jetson-nvmap-gpu and sandbox-rlimits-connect are skipped unless selected)
Summary: 67 passed, 4 failed, 0 cancelled, 2 skipped

Job Result
agent-turn-latency ✅ success
bedrock-runtime-compatible-anthropic ✅ success
brave-search ✅ success
cloud-inference ✅ success
cloud-onboard ✅ success
common-egress-agent ✅ success
concurrent-gateway-ports ✅ success
credential-migration ✅ success
credential-sanitization ✅ success
cron-preflight-inference-local ✅ success
device-auth-health ✅ success
diagnostics ✅ success
docs-validation ✅ success
double-onboard ✅ success
full-e2e ✅ success
gateway-drift-preflight ✅ success
gateway-guard-recovery ✅ success
gateway-health-honest ✅ success
generate-matrix ✅ success
gpu-double-onboard ✅ success
gpu-e2e ✅ success
hermes-channels-add-remove ❌ failure
hermes-channels-stop-start ❌ failure
hermes-dashboard ✅ success
hermes-e2e ✅ success
hermes-inference-switch ✅ success
hermes-root-entrypoint-smoke ✅ success
hermes-sandbox-secret-boundary ✅ success
inference-routing ✅ success
issue-2478-crash-loop-recovery ✅ success
issue-4434-tui-unreachable-inference ✅ success
issue-4462-scope-upgrade-approval ✅ success
jetson-nvmap-gpu ⏭️ skipped
kimi-inference-compat ✅ success
launchable-smoke ✅ success
live ✅ success
messaging-compatible-endpoint ✅ success
model-router-provider-routed-inference ✅ success
network-policy ✅ success
ollama-auth-proxy ✅ success
onboard-negative-paths ✅ success
onboard-repair ✅ success
onboard-resume ✅ success
openclaw-channels-add-remove ❌ failure
openclaw-channels-conflict-guard ✅ success
openclaw-channels-credential-rewrite ✅ success
openclaw-channels-pairing ✅ success
openclaw-channels-stop-start ❌ failure
openclaw-channels-telegram-injection-safety ✅ success
openclaw-channels-token-rotation ✅ success
openclaw-inference-switch ✅ success
openclaw-skill-cli ✅ success
openclaw-tui-chat-correlation ✅ success
openshell-gateway-upgrade ✅ success
openshell-version-pin ✅ success
overlayfs-autofix ✅ success
rebuild-hermes ✅ success
rebuild-hermes-stale-base ✅ success
rebuild-openclaw ✅ success
runtime-overrides ✅ success
sandbox-operations ✅ success
sandbox-rebuild ✅ success
sandbox-rlimits-connect ⏭️ skipped
sandbox-survival ✅ success
security-posture ✅ success
sessions-agents-cli ✅ success
shields-config ✅ success
skill-agent ✅ success
snapshot-commands ✅ success
spark-install ✅ success
state-backup-restore ✅ success
tunnel-lifecycle ✅ success
upgrade-stale-sandbox ✅ success

Explicit-only jobs skipped: sandbox-rlimits-connect (default dispatch excludes the destructive rlimit fork/connect probe unless selected; validate with jobs=sandbox-rlimits-connect or targets=sandbox-rlimits-connect), jetson-nvmap-gpu (default dispatch excludes Jetson until a stable Jetson runner is available; validate with jobs=jetson-nvmap-gpu or targets=jetson-nvmap-gpu).

Failed jobs: hermes-channels-add-remove, hermes-channels-stop-start, openclaw-channels-add-remove, openclaw-channels-stop-start. Check run artifacts for logs.

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 5

Caution

Some comments are outside the diff and can’t be posted inline due to platform limitations.

⚠️ Outside diff range comments (2)
test/e2e/support/e2e-workflow.test.ts (1)

1207-1257: 🔒 Security & Privacy | 🟠 Major | ⚡ Quick win

Restore the Telegram fake-token assertion in this drift case.

.github/workflows/e2e.yaml still hardcodes a fake TELEGRAM_BOT_TOKEN for openclaw-channels-stop-start, but this boundary test now only proves Discord and Teams token enforcement. If validateE2eWorkflowBoundary() stops checking the Telegram token, this suite still passes and a real Telegram credential can slip into the job unnoticed. As per path instructions, **/*.test.{ts,js,mts,mjs,cts,cjs}: Review tests for behavioral confidence rather than implementation lock-in.

🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@test/e2e/support/e2e-workflow.test.ts` around lines 1207 - 1257, The boundary
test for the openclaw-channels-stop-start workflow no longer verifies the fake
Telegram credential, so add back an assertion that validateE2eWorkflowBoundary()
reports the expected TELEGRAM_BOT_TOKEN check alongside the existing token
validations. Update the expectations in e2e-workflow.test.ts around the
openclaw-channels-stop-start case to ensure the job still rejects real Telegram
credentials, using the existing runStep and validateE2eWorkflowBoundary symbols
to locate the drift case.

Source: Path instructions

.github/workflows/regression-e2e.yaml (1)

323-354: 🎯 Functional Correctness | 🟠 Major | ⚡ Quick win

Keep this regression lane scoped to the WhatsApp QR guard.

The job is documented as a hermetic WhatsApp compact-QR regression, but it now runs test/e2e/live/openclaw-channels-pairing.test.ts, which this PR’s context describes as the consolidated Discord/Slack/WhatsApp pairing suite. That broadens a focused regression lane into a multi-channel live suite and can pull in unrelated failures or setup requirements. Point this job back to a dedicated WhatsApp QR entrypoint, or at least add a focused Vitest filter for the WhatsApp case.

🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In @.github/workflows/regression-e2e.yaml around lines 323 - 354, The regression
job is using the consolidated openclaw-channels-pairing.test.ts suite instead of
staying scoped to the WhatsApp QR guard. Update the
openclaw-channels-pairing-whatsapp-qr-e2e workflow step to run only the WhatsApp
pairing path, either by pointing the npx vitest run command at a dedicated
WhatsApp QR test entrypoint or by adding a Vitest filter that targets the
WhatsApp case in openclaw-channels-pairing.test.ts.
🧹 Nitpick comments (1)
test/e2e/live/openclaw-channels-pairing-discord.ts (1)

97-116: 📐 Maintainability & Code Quality | 🔵 Trivial | ⚡ Quick win

Avoid coupling this helper to the private sandbox config layout.

This block shells into /sandbox/.openclaw/openclaw.json and asserts nested channels.discord.accounts.default keys directly. That makes the pairing suite fail on internal config reshuffles even when the public Discord behavior is still correct. The provider lookup, fake-gateway proof, and pairing approval below already cover the observable contract.

🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@test/e2e/live/openclaw-channels-pairing-discord.ts` around lines 97 - 116,
This Discord pairing test is too tightly coupled to the internal sandbox config
structure by reading /sandbox/.openclaw/openclaw.json and asserting nested
channels.discord.accounts.default fields directly. Remove that helper block from
openclaw-channels-pairing-discord.ts and rely on the existing observable checks
in the pairing flow, such as the provider lookup, fake-gateway proof, and
pairing approval assertions, so the test validates behavior instead of private
config layout.
🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

Inline comments:
In `@test/e2e/live/openclaw-channels-conflict-guard.test.ts`:
- Around line 126-136: The commandEnv() helper is pulling in unrelated messaging
credentials from process.env, which can make the
openclaw-channels-conflict-guard test depend on CI/local state instead of only
the Telegram duplicate being exercised. Update commandEnv() to build a minimal
child environment and explicitly remove or omit unrelated messaging variables
such as DISCORD_*, SLACK_*, and any other messaging creds that could affect this
single-channel conflict case, while keeping the Telegram values needed by the
test.

In `@test/e2e/support/e2e-workflow.test.ts`:
- Around line 601-630: The add/remove workflow coverage is split across selector
modes, so each free-standing job should be asserted through both `targets` and
`jobs` in `evaluateE2eWorkflowDispatchSelectors`. Update the contract tests in
`e2e-workflow.test.ts` to mirror the existing loop used for
`gateway-health-honest`, `concurrent-gateway-ports`, and
`openclaw-channels-conflict-guard`, and add both selector assertions for
`openclaw-channels-add-remove` and `hermes-channels-add-remove` to catch typos
in either branch.

In `@test/e2e/support/openclaw-channels-pairing-workflow-boundary.test.ts`:
- Around line 55-60: Update the boundary test in
openclaw-channels-pairing-workflow-boundary to cover the Slack bot token secret
drift as well: in the live step setup and matching assertions, mutate and verify
SLACK_BOT_TOKEN alongside DISCORD_BOT_TOKEN and SLACK_APP_TOKEN. Use the
existing pairingJob.steps lookup for "Run OpenClaw channels pairing live tests"
and the installOpenShell-related assertions so the test fails if SLACK_BOT_TOKEN
is no longer enforced.

In `@test/regression-e2e-workflow.test.ts`:
- Around line 59-63: The regression test is only asserting absence of the env
var in the generated shell command text, which can miss changes to the workflow
contract. Update test/regression-e2e-workflow.test.ts to inspect the parsed
workflow objects for the relevant job/step env maps, using the job and step
structures around the existing runText assertions. Keep the existing checks for
the e2e-live command and target test path, but replace the NEMOCLAW_RUN_LIVE_E2E
source-text assertion with a direct expectation on job.env or step.env so the
boundary is validated behaviorally.

In `@tools/e2e/workflow-boundary.mts`:
- Around line 4272-4289: The step validation in workflow-boundary should also
block unexpected GITHUB_TOKEN exposure for these job steps, since the current
checks in the step loop only cover NVIDIA_INFERENCE_API_KEY and Docker Hub
secrets. Update the same validation logic around asSteps/job.steps and the
per-step loop to call the existing secret-exposure guard for GITHUB_TOKEN on
every step except any explicitly allowed case, matching the neighboring
validators’ coverage so credential-rewrite and telegram-injection-safety jobs
cannot pass with `${{ github.token }}` in env or run content.

---

Outside diff comments:
In @.github/workflows/regression-e2e.yaml:
- Around line 323-354: The regression job is using the consolidated
openclaw-channels-pairing.test.ts suite instead of staying scoped to the
WhatsApp QR guard. Update the openclaw-channels-pairing-whatsapp-qr-e2e workflow
step to run only the WhatsApp pairing path, either by pointing the npx vitest
run command at a dedicated WhatsApp QR test entrypoint or by adding a Vitest
filter that targets the WhatsApp case in openclaw-channels-pairing.test.ts.

In `@test/e2e/support/e2e-workflow.test.ts`:
- Around line 1207-1257: The boundary test for the openclaw-channels-stop-start
workflow no longer verifies the fake Telegram credential, so add back an
assertion that validateE2eWorkflowBoundary() reports the expected
TELEGRAM_BOT_TOKEN check alongside the existing token validations. Update the
expectations in e2e-workflow.test.ts around the openclaw-channels-stop-start
case to ensure the job still rejects real Telegram credentials, using the
existing runStep and validateE2eWorkflowBoundary symbols to locate the drift
case.

---

Nitpick comments:
In `@test/e2e/live/openclaw-channels-pairing-discord.ts`:
- Around line 97-116: This Discord pairing test is too tightly coupled to the
internal sandbox config structure by reading /sandbox/.openclaw/openclaw.json
and asserting nested channels.discord.accounts.default fields directly. Remove
that helper block from openclaw-channels-pairing-discord.ts and rely on the
existing observable checks in the pairing flow, such as the provider lookup,
fake-gateway proof, and pairing approval assertions, so the test validates
behavior instead of private config layout.
🪄 Autofix (Beta)

Fix all unresolved CodeRabbit comments on this PR:

  • Push a commit to this branch (recommended)
  • Create a new PR with the fixes

ℹ️ Review info
⚙️ Run configuration

Configuration used: Path: .coderabbit.yaml

Review profile: CHILL

Plan: Enterprise

Run ID: 03b9ebae-9958-4cb1-872d-3529df518151

📥 Commits

Reviewing files that changed from the base of the PR and between 79d9cbe and 7ac3c53.

📒 Files selected for processing (42)
  • .github/workflows/brev-nightly-e2e.yaml
  • .github/workflows/e2e-branch-validation.yaml
  • .github/workflows/e2e.yaml
  • .github/workflows/regression-e2e.yaml
  • docs/manage-sandboxes/messaging-channels.mdx
  • docs/reference/troubleshooting.mdx
  • src/lib/actions/sandbox/policy-explain.ts
  • test/e2e-advisor-targets.test.ts
  • test/e2e-release-gate-workflow.test.ts
  • test/e2e/brev-e2e.test.ts
  • test/e2e/live/channels-add-remove.test.ts
  • test/e2e/live/channels-lifecycle-helpers.ts
  • test/e2e/live/channels-stop-start-helpers.ts
  • test/e2e/live/channels-stop-start-safety.ts
  • test/e2e/live/channels-stop-start.test.ts
  • test/e2e/live/hermes-channels-add-remove.test.ts
  • test/e2e/live/hermes-channels-stop-start.test.ts
  • test/e2e/live/hermes-discord.test.ts
  • test/e2e/live/hermes-slack-e2e-helpers.ts
  • test/e2e/live/hermes-slack-e2e.test.ts
  • test/e2e/live/openclaw-channels-add-remove.test.ts
  • test/e2e/live/openclaw-channels-conflict-guard.test.ts
  • test/e2e/live/openclaw-channels-credential-rewrite-helpers.ts
  • test/e2e/live/openclaw-channels-credential-rewrite.test.ts
  • test/e2e/live/openclaw-channels-pairing-discord.ts
  • test/e2e/live/openclaw-channels-pairing-slack.ts
  • test/e2e/live/openclaw-channels-pairing-whatsapp-qr.ts
  • test/e2e/live/openclaw-channels-pairing.test.ts
  • test/e2e/live/openclaw-channels-stop-start.test.ts
  • test/e2e/live/openclaw-channels-telegram-injection-safety.test.ts
  • test/e2e/live/openclaw-channels-token-rotation.test.ts
  • test/e2e/live/openclaw-discord-pairing.test.ts
  • test/e2e/live/openclaw-pairing-helpers.ts
  • test/e2e/live/openclaw-slack-pairing.test.ts
  • test/e2e/live/phase6-messaging-helpers.ts
  • test/e2e/live/rebuild-openclaw.test.ts
  • test/e2e/support/e2e-workflow.test.ts
  • test/e2e/support/openclaw-channels-pairing-workflow-boundary.test.ts
  • test/e2e/support/openclaw-discord-workflow-boundary.test.ts
  • test/regression-e2e-workflow.test.ts
  • test/whatsapp-qr-compact.test.ts
  • tools/e2e/workflow-boundary.mts
💤 Files with no reviewable changes (10)
  • test/e2e/live/openclaw-slack-pairing.test.ts
  • test/e2e/live/openclaw-discord-pairing.test.ts
  • test/e2e/support/openclaw-discord-workflow-boundary.test.ts
  • test/e2e/live/channels-stop-start.test.ts
  • test/e2e/live/hermes-slack-e2e.test.ts
  • test/e2e/live/channels-stop-start-safety.ts
  • test/e2e/live/hermes-slack-e2e-helpers.ts
  • test/e2e/live/channels-add-remove.test.ts
  • test/e2e/live/hermes-discord.test.ts
  • test/e2e/live/channels-stop-start-helpers.ts

Comment thread test/e2e/live/openclaw-channels-conflict-guard.test.ts Outdated
Comment thread test/e2e/support/e2e-workflow.test.ts Outdated
Comment thread test/regression-e2e-workflow.test.ts Outdated
Comment thread tools/e2e/workflow-boundary.mts
@github-actions

Copy link
Copy Markdown
Contributor

Vitest E2E Target Results — ❌ Some jobs failed

Run: 28445695305
Workflow ref: test/e2e-messaging-refactor-6022
Requested targets: openclaw-channels-add-remove,hermes-channels-add-remove,openclaw-channels-stop-start,hermes-channels-stop-start
Requested jobs: (default — all default-enabled free-standing jobs; explicit-only jobs such as jetson-nvmap-gpu and sandbox-rlimits-connect are skipped unless selected)
Summary: 2 passed, 2 failed, 0 cancelled, 0 skipped

Job Result
hermes-channels-add-remove ❌ failure
hermes-channels-stop-start ✅ success
openclaw-channels-add-remove ❌ failure
openclaw-channels-stop-start ✅ success

Failed jobs: hermes-channels-add-remove, openclaw-channels-add-remove. Check run artifacts for logs.

Signed-off-by: San Dang <sdang@nvidia.com>
@github-actions

Copy link
Copy Markdown
Contributor

Vitest E2E Target Results — ⚠️ Some jobs cancelled — partial pass

Run: 28450217578
Workflow ref: test/e2e-messaging-refactor-6022
Requested targets: openclaw-channels-add-remove,hermes-channels-add-remove
Requested jobs: (default — all default-enabled free-standing jobs; explicit-only jobs such as jetson-nvmap-gpu and sandbox-rlimits-connect are skipped unless selected)
Summary: 1 passed, 0 failed, 1 cancelled, 0 skipped

Job Result
hermes-channels-add-remove ⚠️ cancelled
openclaw-channels-add-remove ✅ success

sandl99 added 3 commits June 30, 2026 21:55
Signed-off-by: San Dang <sdang@nvidia.com>
Signed-off-by: San Dang <sdang@nvidia.com>
Signed-off-by: San Dang <sdang@nvidia.com>
@sandl99 sandl99 added area: messaging Messaging channels, bridges, manifests, or channel lifecycle area: e2e End-to-end tests, nightly failures, or validation infrastructure v0.0.71 and removed area: docs Documentation, examples, guides, or docs build labels Jun 30, 2026
@github-actions

Copy link
Copy Markdown
Contributor

Vitest E2E Target Results — ❌ Some jobs failed

Run: 28455933667
Workflow ref: test/e2e-messaging-refactor-6022
Requested targets: (default — all supported)
Requested jobs: openclaw-channels-credential-rewrite,openclaw-channels-conflict-guard,openclaw-channels-add-remove,hermes-channels-add-remove,openclaw-channels-stop-start,hermes-channels-stop-start,openclaw-channels-token-rotation,openclaw-channels-pairing,openclaw-channels-telegram-injection-safety,network-policy
Summary: 9 passed, 1 failed, 0 cancelled, 0 skipped

Job Result
hermes-channels-add-remove ❌ failure
hermes-channels-stop-start ✅ success
network-policy ✅ success
openclaw-channels-add-remove ✅ success
openclaw-channels-conflict-guard ✅ success
openclaw-channels-credential-rewrite ✅ success
openclaw-channels-pairing ✅ success
openclaw-channels-stop-start ✅ success
openclaw-channels-telegram-injection-safety ✅ success
openclaw-channels-token-rotation ✅ success

Failed jobs: hermes-channels-add-remove. Check run artifacts for logs.

@github-actions

Copy link
Copy Markdown
Contributor

Vitest E2E Target Results — ❌ Some jobs failed

Run: 28459843100
Workflow ref: test/e2e-messaging-refactor-6022
Requested targets: (default — all supported)
Requested jobs: hermes-channels-add-remove
Summary: 0 passed, 1 failed, 0 cancelled, 0 skipped

Job Result
hermes-channels-add-remove ❌ failure

Failed jobs: hermes-channels-add-remove. Check run artifacts for logs.

@prekshivyas prekshivyas left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pure rename/restructure — no assertions dropped, no if-statements in test bodies, add-remove jobs expand channel coverage. LGTM.

@sandl99 sandl99 removed the v0.0.71 label Jun 30, 2026
Signed-off-by: San Dang <sdang@nvidia.com>
Signed-off-by: San Dang <sdang@nvidia.com>
@github-actions

github-actions Bot commented Jul 1, 2026

Copy link
Copy Markdown
Contributor

Vitest E2E Target Results — ✅ All requested jobs passed

Run: 28491284015
Workflow ref: test/e2e-messaging-refactor-6022
Requested targets: (default — all supported)
Requested jobs: hermes-channels-add-remove
Summary: 1 passed, 0 failed, 0 cancelled, 0 skipped

Job Result
hermes-channels-add-remove ✅ success

@github-actions

github-actions Bot commented Jul 1, 2026

Copy link
Copy Markdown
Contributor

Vitest E2E Target Results — ✅ All requested jobs passed

Run: 28492704625
Workflow ref: test/e2e-messaging-refactor-6022
Requested targets: (default — all supported)
Requested jobs: openclaw-channels-credential-rewrite,openclaw-channels-conflict-guard,openclaw-channels-add-remove,hermes-channels-add-remove,openclaw-channels-stop-start,hermes-channels-stop-start,openclaw-channels-token-rotation,openclaw-channels-pairing,openclaw-channels-telegram-injection-safety,network-policy
Summary: 10 passed, 0 failed, 0 cancelled, 0 skipped

Job Result
hermes-channels-add-remove ✅ success
hermes-channels-stop-start ✅ success
network-policy ✅ success
openclaw-channels-add-remove ✅ success
openclaw-channels-conflict-guard ✅ success
openclaw-channels-credential-rewrite ✅ success
openclaw-channels-pairing ✅ success
openclaw-channels-stop-start ✅ success
openclaw-channels-telegram-injection-safety ✅ success
openclaw-channels-token-rotation ✅ success

@github-actions github-actions Bot mentioned this pull request Jul 1, 2026
10 of 21 tasks
@github-actions

github-actions Bot commented Jul 1, 2026

Copy link
Copy Markdown
Contributor

Vitest E2E Target Results — ✅ All requested jobs passed

Run: 28508508090
Workflow ref: test/e2e-messaging-refactor-6022
Requested targets: (default — all supported)
Requested jobs: openclaw-channels-credential-rewrite,openclaw-channels-conflict-guard,openclaw-channels-add-remove,hermes-channels-add-remove,openclaw-channels-stop-start,hermes-channels-stop-start,openclaw-channels-token-rotation,openclaw-channels-pairing,openclaw-channels-telegram-injection-safety,network-policy
Summary: 10 passed, 0 failed, 0 cancelled, 0 skipped

Job Result
hermes-channels-add-remove ✅ success
hermes-channels-stop-start ✅ success
network-policy ✅ success
openclaw-channels-add-remove ✅ success
openclaw-channels-conflict-guard ✅ success
openclaw-channels-credential-rewrite ✅ success
openclaw-channels-pairing ✅ success
openclaw-channels-stop-start ✅ success
openclaw-channels-telegram-injection-safety ✅ success
openclaw-channels-token-rotation ✅ success

@github-actions

github-actions Bot commented Jul 1, 2026

Copy link
Copy Markdown
Contributor

Vitest E2E Target Results — ⚠️ Some jobs cancelled — partial pass

Run: 28510597888
Workflow ref: test/e2e-messaging-refactor-6022
Requested targets: (default — all supported)
Requested jobs: openclaw-channels-credential-rewrite,openclaw-channels-conflict-guard,openclaw-channels-add-remove,hermes-channels-add-remove,openclaw-channels-stop-start,hermes-channels-stop-start,openclaw-channels-token-rotation,openclaw-channels-pairing,openclaw-channels-telegram-injection-safety,network-policy
Summary: 1 passed, 0 failed, 9 cancelled, 0 skipped

Job Result
hermes-channels-add-remove ⚠️ cancelled
hermes-channels-stop-start ⚠️ cancelled
network-policy ⚠️ cancelled
openclaw-channels-add-remove ⚠️ cancelled
openclaw-channels-conflict-guard ✅ success
openclaw-channels-credential-rewrite ⚠️ cancelled
openclaw-channels-pairing ⚠️ cancelled
openclaw-channels-stop-start ⚠️ cancelled
openclaw-channels-telegram-injection-safety ⚠️ cancelled
openclaw-channels-token-rotation ⚠️ cancelled

@github-actions

github-actions Bot commented Jul 1, 2026

Copy link
Copy Markdown
Contributor

Vitest E2E Target Results — ✅ All requested jobs passed

Run: 28510840868
Workflow ref: test/e2e-messaging-refactor-6022
Requested targets: (default — all supported)
Requested jobs: openclaw-channels-credential-rewrite,openclaw-channels-conflict-guard,openclaw-channels-add-remove,hermes-channels-add-remove,openclaw-channels-stop-start,hermes-channels-stop-start,openclaw-channels-token-rotation,openclaw-channels-pairing,openclaw-channels-telegram-injection-safety,network-policy
Summary: 10 passed, 0 failed, 0 cancelled, 0 skipped

Job Result
hermes-channels-add-remove ✅ success
hermes-channels-stop-start ✅ success
network-policy ✅ success
openclaw-channels-add-remove ✅ success
openclaw-channels-conflict-guard ✅ success
openclaw-channels-credential-rewrite ✅ success
openclaw-channels-pairing ✅ success
openclaw-channels-stop-start ✅ success
openclaw-channels-telegram-injection-safety ✅ success
openclaw-channels-token-rotation ✅ success

@github-actions

github-actions Bot commented Jul 1, 2026

Copy link
Copy Markdown
Contributor

Vitest E2E Target Results — ⚠️ Some jobs cancelled — partial pass

Run: 28514144167
Workflow ref: test/e2e-messaging-refactor-6022
Requested targets: (default — all supported)
Requested jobs: openclaw-channels-credential-rewrite,openclaw-channels-conflict-guard,openclaw-channels-add-remove,hermes-channels-add-remove,openclaw-channels-stop-start,hermes-channels-stop-start,openclaw-channels-token-rotation,openclaw-channels-pairing,openclaw-channels-telegram-injection-safety,network-policy
Summary: 2 passed, 0 failed, 8 cancelled, 0 skipped

Job Result
hermes-channels-add-remove ⚠️ cancelled
hermes-channels-stop-start ⚠️ cancelled
network-policy ⚠️ cancelled
openclaw-channels-add-remove ⚠️ cancelled
openclaw-channels-conflict-guard ✅ success
openclaw-channels-credential-rewrite ⚠️ cancelled
openclaw-channels-pairing ⚠️ cancelled
openclaw-channels-stop-start ⚠️ cancelled
openclaw-channels-telegram-injection-safety ✅ success
openclaw-channels-token-rotation ⚠️ cancelled

@github-actions

github-actions Bot commented Jul 1, 2026

Copy link
Copy Markdown
Contributor

Vitest E2E Target Results — ⚠️ Some jobs cancelled — partial pass

Run: 28514531378
Workflow ref: test/e2e-messaging-refactor-6022
Requested targets: (default — all supported)
Requested jobs: openclaw-channels-credential-rewrite,openclaw-channels-conflict-guard,openclaw-channels-add-remove,hermes-channels-add-remove,openclaw-channels-stop-start,hermes-channels-stop-start,openclaw-channels-token-rotation,openclaw-channels-pairing,openclaw-channels-telegram-injection-safety,network-policy
Summary: 1 passed, 0 failed, 9 cancelled, 0 skipped

Job Result
hermes-channels-add-remove ⚠️ cancelled
hermes-channels-stop-start ⚠️ cancelled
network-policy ⚠️ cancelled
openclaw-channels-add-remove ⚠️ cancelled
openclaw-channels-conflict-guard ✅ success
openclaw-channels-credential-rewrite ⚠️ cancelled
openclaw-channels-pairing ⚠️ cancelled
openclaw-channels-stop-start ⚠️ cancelled
openclaw-channels-telegram-injection-safety ⚠️ cancelled
openclaw-channels-token-rotation ⚠️ cancelled

@github-actions

github-actions Bot commented Jul 1, 2026

Copy link
Copy Markdown
Contributor

Vitest E2E Target Results — ❌ Some jobs failed

Run: 28514977239
Workflow ref: test/e2e-messaging-refactor-6022
Requested targets: (default — all supported)
Requested jobs: messaging-compatible-endpoint,openclaw-channels-token-rotation,openclaw-channels-credential-rewrite,openclaw-channels-conflict-guard,openclaw-channels-add-remove,hermes-channels-add-remove,openclaw-channels-telegram-injection-safety,openclaw-channels-stop-start,hermes-channels-stop-start,openclaw-channels-pairing
Summary: 9 passed, 1 failed, 0 cancelled, 0 skipped

Job Result
hermes-channels-add-remove ✅ success
hermes-channels-stop-start ✅ success
messaging-compatible-endpoint ✅ success
openclaw-channels-add-remove ✅ success
openclaw-channels-conflict-guard ✅ success
openclaw-channels-credential-rewrite ❌ failure
openclaw-channels-pairing ✅ success
openclaw-channels-stop-start ✅ success
openclaw-channels-telegram-injection-safety ✅ success
openclaw-channels-token-rotation ✅ success

Failed jobs: openclaw-channels-credential-rewrite. Check run artifacts for logs.

ericksoa added a commit that referenced this pull request Jul 2, 2026
## Summary

Routes ordinary native CDI Linux through OpenShell 0.0.71's native
`--gpu` path instead of the legacy Docker container swap. The legacy
path remains available for WSL, Jetson, and explicit
`NEMOCLAW_DOCKER_GPU_PATCH=1`, with its OpenShell supervisor command
boundary and rollback diagnostics hardened.

## Related Issue

Related to #6110

## Changes

- Use native OpenShell GPU injection by default on ordinary native CDI
Linux.
- Document the `NEMOCLAW_DOCKER_GPU_PATCH` auto, forced-legacy, and
native-routing behavior for ordinary native Linux, Docker Desktop WSL,
and Jetson/Tegra.
- Keep `NEMOCLAW_DOCKER_GPU_PATCH=1` as the explicit legacy-swap force
control and `=0` as the existing native opt-out; Docker Desktop WSL
still ignores `=0`, and Jetson keeps its compatibility default.
- Preserve OpenShell's supervisor entrypoint on the legacy swap: Docker
receives no command tail, while the workload stays in
`OPENSHELL_SANDBOX_COMMAND`.
- Validate legacy startup tokens before stopping or renaming the
original container and serialize extra-placeholder keys as one
comma-delimited token.
- Defer every legacy recreate through the same supervisor-wait/finalize
boundary so failed clones are captured before rollback on both create
timing paths.
- Persist only allowlisted/redacted failed-clone topology, state,
process, network, and log evidence, with a 10-second total / 2-second
per-call budget so diagnostics cannot materially delay rollback.
- Permit only the canonical OpenShell Docker/Podman TLS key path in the
Hermes runtime environment; arbitrary values and persisted `.env`
entries remain rejected.
- Refresh the Dockerfile integrity pin for the changed validator so
production Hermes images fail closed on any later digest drift.
- Prove native and legacy Docker command boundaries separately,
including Ready/CUDA status, supervisor PID 1, placeholder transport,
config hashes, no backup-container leak, and inference requests.

## Type of Change

- [ ] Code change (feature, bug fix, or refactor)
- [x] Code change with doc updates
- [ ] Doc only (prose changes, no code sample modifications)
- [ ] Doc only (includes code sample changes)

## Quality Gates

- [x] Tests added or updated for changed behavior
- [ ] Existing tests cover changed behavior — justification:
- [ ] Tests not applicable — justification:
- [x] Docs updated for user-facing behavior changes
- [ ] Docs not applicable — justification:
- [x] Sensitive paths changed (security, policy, credentials, preflight,
onboarding, inference, runner, sandbox, or messaging)
- [x] Sensitive-path review completed or maintainer-approved waiver
recorded — reviewer/approval link/justification: independent exact-diff
review found no remaining substantive issue after startup-envelope
secret redaction, bounded pre-rollback capture, unified finalize
ordering, and route-specific live assertions
- [ ] Non-success, skipped, or missing CI check accepted by maintainer —
check name, approval link, and follow-up issue:

## Verification

- [x] PR description includes the DCO sign-off declaration and every
commit appears as `Verified` in GitHub
- [ ] Git hooks passed during commit and push, or `npx prek run
--from-ref main --to-ref HEAD` passes
- [x] Targeted tests pass for changed behavior
- [ ] Full `npm test` passes (broad runtime changes only)
- [x] Quality Gates section completed with required justifications or
waivers
- [x] No secrets, API keys, or credentials committed
- [ ] `npm run docs` builds without warnings (doc changes only)
- [x] Doc pages follow the [style
guide](https://github.com/NVIDIA/NemoClaw/blob/main/docs/CONTRIBUTING.md)
(doc changes only)
- [ ] New doc pages include SPDX header and frontmatter (new pages only)

Local exact-candidate verification at
`d76f1647a5f354ba04737a6a049b82bfbf6d5454`:

- CLI build and typecheck passed.
- Hermes GPU support/client/workflow coverage passed 48/48; the built
gateway-cleanup module resolves through Node, and runtime cleanup,
registration removal, and bind availability remain fail-closed.
- The shared Docker GPU diagnostic collector owns redaction for every
text/JSON artifact and returned summary; direct conventional `*_KEY` and
custom-placeholder canaries, JSON-validity, inspect-before-write,
exhausted-budget, and collector-owned top regressions pass.
- All 12 Docker GPU suites pass 126/126; the exact-head focused
E2E-support set passes 39/39, including six process-token self-match
regressions, the scrubbed integrity-proof environment, and total
forbidden-marker count.
- Conditional scan, source-shape, test-size, Biome, repository checks,
commit hooks, and push typecheck passed for the final harness
correction; only the documented macOS-invalid full CLI hook lane was
excluded.
- The forbidden-marker request sensor support suite passed 14/14 and
records counts only, never raw request bodies or marker values.
- OpenShell transport-boundary coverage passed 4/4; Docker GPU
command-envelope coverage passed 6/6; extra-placeholder parsing coverage
passed 16/16.
- Hermes validator and wrapper integrity pins match their source SHA256
digests; hadolint and diff checks pass.
- Hermes startup/boundary coverage passed 43/43 locally; the Linux-only
wrapper cases are delegated to exact-head CI.
- The full local CLI hook is not a valid gate on this macOS Node 22
host: unchanged `a1fc52c7` TypeScript entrypoints fail through the
CommonJS preload with `ERR_UNKNOWN_FILE_EXTENSION`; exact-head Linux CI
remains required.
- Hermes runtime-guard plus current-main docs regression tests: 24/24
passed.
- Hermes workflow-boundary test passed.
- Project-boundary, project-membership, test-title, source-shape,
test-size, targeted Biome, and diff checks passed.
- `npm run docs:sync-agent-variants`, `npm run
docs:check-agent-variants`, and `npm run docs` pass; Fern reports 0
errors and 2 existing warnings.

Live A/B evidence:

- Forced legacy swap at `97e3e7e1`: [run
28554699811](https://github.com/NVIDIA/NemoClaw/actions/runs/28554699811)
reproduced sustained OpenShell `Error` followed by safe rollback. It
also exposed and now closes a lifecycle instrumentation gap: the
post-create `ensureApplied()` branch bypassed pre-rollback capture.
- Native OpenShell route at signed diagnostic SHA `15f50182`: [run
28555110558](https://github.com/NVIDIA/NemoClaw/actions/runs/28555110558)
completed onboarding with exit 0, reached `Phase: Ready`, reported `CUDA
verified`, and sent authenticated Hermes chat-completions requests to
the hermetic inference endpoint. The job's only failure was the
now-fixed test regex not stripping ANSI around `Ready`.
- Prior native evidence at `7cb219d9`: [full run
28559814959](https://github.com/NVIDIA/NemoClaw/actions/runs/28559814959)
and [second pass
28559816026](https://github.com/NVIDIA/NemoClaw/actions/runs/28559816026)
both reached Ready/CUDA with clean runtime and teardown; their Hermes
proof stopped on the now-fixed ANSI matcher before downstream
assertions.
- Prior forced-legacy diagnostic on production parent `a1fc52c7` plus
workflow-only child `6d0cf6a5`: [run
28561607207](https://github.com/NVIDIA/NemoClaw/actions/runs/28561607207)
selected the legacy swap and rolled back cleanly, but failed because the
Hermes boundary rejected the driver-owned OpenShell `OPENSHELL_TLS_KEY`
path. Candidate `54cf259d` adds an exact runtime-only allowance with
negative boundary tests.
- Prior exact-head set at `a1fc52c7`: [run
28561548749](https://github.com/NVIDIA/NemoClaw/actions/runs/28561548749)
proved the GPU and security companion jobs, while Hermes GPU stopped at
a sandbox-user `/proc` permission probe after Ready/CUDA. Candidate
`54cf259d` keeps the same proof but runs it as root and restricts the
match to the exact `nemoclaw-start` process.
- Prior second native pass at `a1fc52c7`: [run
28561555945](https://github.com/NVIDIA/NemoClaw/actions/runs/28561555945)
reproduced only the same harness permission failure after Ready/CUDA,
correct PID 1 topology, authenticated inference, zero forbidden-marker
matches, and clean teardown.
- Superseded six-job set at `54cf259d`: [run
28564960504](https://github.com/NVIDIA/NemoClaw/actions/runs/28564960504)
exposed the stale Dockerfile validator digest and was canceled before
runtime proof. Candidate `970803a4` updates the integrity pin, retained
by final head `c5a67c4c`.
- Superseded second native pass at `54cf259d`: [run
28564973806](https://github.com/NVIDIA/NemoClaw/actions/runs/28564973806)
was canceled during pre-cleanup and supplies no acceptance evidence.
- Superseded forced-legacy proof on production parent `54cf259d` plus
child `69f4e1b2`: [run
28564983760](https://github.com/NVIDIA/NemoClaw/actions/runs/28564983760)
was canceled during pre-cleanup and supplies no acceptance evidence.
- Superseded six-job set at `970803a4`: [run
28565197328](https://github.com/NVIDIA/NemoClaw/actions/runs/28565197328)
was canceled before acceptance execution when the canonical
placeholder-format advisor fix advanced the head.
- Superseded second native pass at `970803a4`: [run
28565207911](https://github.com/NVIDIA/NemoClaw/actions/runs/28565207911)
was canceled before runner assignment and supplies no acceptance
evidence.
- Superseded forced-legacy proof on production parent `970803a4` plus
child `b4d5679e`: [run
28565222881](https://github.com/NVIDIA/NemoClaw/actions/runs/28565222881)
was canceled before runner assignment and supplies no acceptance
evidence.
- Superseded six-job set at `7335903b`: [run
28565576066](https://github.com/NVIDIA/NemoClaw/actions/runs/28565576066)
was intentionally canceled when the documentation gap advanced the
candidate; the root-entrypoint smoke passed, but the remaining lanes
provide no complete acceptance proof.
- Superseded second native pass at `7335903b`: [run
28565587094](https://github.com/NVIDIA/NemoClaw/actions/runs/28565587094)
was canceled before acceptance execution and supplies no acceptance
evidence.
- Superseded forced-legacy proof on production parent `7335903b` plus
child `091e16fd`: [run
28565603460](https://github.com/NVIDIA/NemoClaw/actions/runs/28565603460)
was canceled before acceptance execution and supplies no acceptance
evidence.
- Superseded six-job set at `c5a67c4c`: [run
28566083673](https://github.com/NVIDIA/NemoClaw/actions/runs/28566083673)
reached the native GPU/runtime proofs before the obsolete raw
strict-hash assertion failed; messaging independently hit the
process-probe self-match fixed by merged #6167. GPU, root-entrypoint,
secret-boundary, and credential companion lanes passed.
- Superseded second native pass at `c5a67c4c`: [run
28566083641](https://github.com/NVIDIA/NemoClaw/actions/runs/28566083641)
proved native routing, Ready/CUDA, `nvidia-smi`, `/proc`, `cuInit(0)=0`,
PID 1, authenticated inference, and cleanup, then failed only the
obsolete raw strict-hash assertion.
- Superseded forced-legacy proof on production parent `c5a67c4c` plus
child `48c46a7f`: [run
28566083589](https://github.com/NVIDIA/NemoClaw/actions/runs/28566083589)
proved the same runtime boundary on the legacy route, then failed only
the obsolete raw strict-hash assertion.
- Superseded six-job set at `a04a70ac`: [run
28568069499](https://github.com/NVIDIA/NemoClaw/actions/runs/28568069499)
exposed a pre-onboarding harness defect: direct Vitest import of the
production cleanup helper could not resolve its lazy CommonJS TypeScript
dependencies. Companion results do not count as final-head evidence.
- Superseded second native pass at `a04a70ac`: [run
28568069558](https://github.com/NVIDIA/NemoClaw/actions/runs/28568069558)
failed at the same pre-onboarding cleanup boundary and supplies no
runtime acceptance evidence.
- Superseded forced-legacy proof on production parent `a04a70ac` plus
child `085a3b7d`: [run
28568069530](https://github.com/NVIDIA/NemoClaw/actions/runs/28568069530)
failed at the same pre-onboarding cleanup boundary and supplies no
runtime acceptance evidence.
- Superseded six-job set at `6ac4ebc8`: [run
28568490864](https://github.com/NVIDIA/NemoClaw/actions/runs/28568490864)
exposed a clean-runner preinstall edge: the compiled cleanup child was
invoked before OpenShell existed and failed before onboarding. Companion
results do not count as final-head evidence.
- Superseded second native pass at `6ac4ebc8`: [run
28568494928](https://github.com/NVIDIA/NemoClaw/actions/runs/28568494928)
failed at the same pre-onboarding cleanup boundary and supplies no
runtime acceptance evidence.
- Superseded forced-legacy proof on production parent `6ac4ebc8` plus
child `5d3742a7`: [run
28568501430](https://github.com/NVIDIA/NemoClaw/actions/runs/28568501430)
failed at the same pre-onboarding cleanup boundary and supplies no
runtime acceptance evidence.
- Superseded six-job set at `65b06d64`: [run
28568954862](https://github.com/NVIDIA/NemoClaw/actions/runs/28568954862)
passed all six jobs, but the candidate advanced to close the
advisor-confirmed shared diagnostic-redaction boundary and two
proof-hardening review threads.
- Superseded second native pass at `65b06d64`: [run
28568959028](https://github.com/NVIDIA/NemoClaw/actions/runs/28568959028)
passed the full native runtime proof but is not final-head evidence.
- Superseded forced-legacy proof on production parent `65b06d64` plus
child `2c6dca1b`: [run
28568966557](https://github.com/NVIDIA/NemoClaw/actions/runs/28568966557)
passed but is not final-parent evidence.
- Final six-job exact-head set at `d76f1647`: [run
28601346031](https://github.com/NVIDIA/NemoClaw/actions/runs/28601346031)
passed all six requested jobs. Native GPU, Hermes startup,
root-entrypoint, secret-boundary, credential-sanitization, and messaging
proofs are green; all 21 messaging raw-token surface probes are
`ABSENT`, and every cleanup record has zero failures.
- Final second native Hermes GPU pass at `d76f1647`: [run
28601348023](https://github.com/NVIDIA/NemoClaw/actions/runs/28601348023)
passed 9/9 assertions. Artifact `8043702467`
(`sha256:0ef929fa2478f5c9579ad2f282c46de8fe4d6eefb04acb062de1af29b4eb002c`)
proves native routing, Ready/CUDA, `nvidia-smi`, `/proc`, successful
`cuInit(0)`, OpenShell PID 1, one container/no backup, authenticated
inference with zero forbidden-marker matches, and clean teardown.
- Final failed-clone rollback proof checks out exact production SHA
`d76f1647` from signed workflow-only child `4c49b5bc`: [run
28602166456](https://github.com/NVIDIA/NemoClaw/actions/runs/28602166456)
passed. Artifact `8044114375`
(`sha256:b5cff9cc7e7cf361f55e074a265c8cfefd3e6dc21ad0d300b140d2a749cde00b`)
records clone exit 137 with `failure_kind=patched_container_failed` and
`rolled_back=no` before finalize, then `rolled_back=yes`, exactly one
running original container, no backup leak, guard-observed clone
removal, clean canary scans, and clean fixture teardown.
- Final forced-legacy success proof on exact production parent
`d76f1647` plus signed workflow-only child `078a372d`: [run
28603335692](https://github.com/NVIDIA/NemoClaw/actions/runs/28603335692)
passed 9/9 assertions. Artifact `8044550700`
(`sha256:56d97c5caa8536ebff735ff600399ac7fafa2fed71206de1f5470e7d15549f6f`)
proves `gpuRoute=legacy-patch`, `--device nvidia.com/gpu=all`,
Ready/CUDA with all three GPU probes, correct OpenShell PID 1/command
envelope, one container/no backup, integrity and negative guard checks,
two authenticated inference requests with zero forbidden-marker matches,
a clean artifact canary scan, and clean teardown.

Source-of-truth review for the retained compatibility path:

- **Invalid state:** the legacy swap temporarily leaves a stopped backup
and running clone with the same OpenShell sandbox ID.
- **Source boundary:** OpenShell's Docker driver reconciles container
summaries into a map keyed only by sandbox ID; 0.0.71 can let the
stopped backup overwrite the running clone and drive the gateway into
terminal `Error`.
- **Source-fix constraint:** the NemoClaw-supported OpenShell release
does not contain deterministic active-container selection. The focused
source fix is open as
[NVIDIA/OpenShell#2116](NVIDIA/OpenShell#2116)
and passes 96/96 Docker-driver tests, strict clippy, and formatting, but
is not yet released or pinned here.
- **Regression coverage:** routing tests pin native auto / forced legacy
/ WSL / Jetson behavior; recreate tests pin capture-before-finalize on
both create timing paths; the secret canary uses the actual single
`OPENSHELL_SANDBOX_COMMAND=env ...` envelope; the live test pins native
and legacy runtime topology separately.
- **Removal condition:** remove the legacy swap and its
rollback/diagnostic modules after WSL and Jetson are proven on native
OpenShell GPU injection and the supported OpenShell floor contains
deterministic duplicate-container reconciliation.
- **WSL boundary:** Docker Desktop WSL does not expose a usable native
CDI route to this flow, so WSL retains the compatibility path and
ignores `NEMOCLAW_DOCKER_GPU_PATCH=0`; routing tests lock that behavior.
Remove it when Docker Desktop exposes usable `nvidia.com/gpu` CDI
devices to the WSL distro.
- **Jetson boundary:** Tegra `/dev/nvmap` and `/dev/nvhost-*` device
ownership requires host group propagation for the non-root sandbox user;
group-add tests lock that behavior. Remove it only when the
platform/runtime supplies equivalent access without the compatibility
recreate.

Diagnostic redaction boundary:

- **Source-of-truth invariant:** `collectDockerGpuPatchDiagnostics()`
constructs the trusted per-bundle redactor, discovers conventional and
custom-placeholder values from every known/discovered full inspect
before writing, recursively redacts JSON values, and publishes every
summary, Docker, OpenShell, and pre-rollback top artifact through that
boundary.
- **Bounded pre-rollback path:** the caller contributes only additive
values discovered from the failed clone before snapshot capture so the
shared 10-second budget cannot hide opaque values; the collector still
performs its own discovery and owns every write. Direct raw-caller and
exhausted-budget regressions scan all artifacts and returned summaries.
- **Removal condition:** remove the additive pre-rollback discovery only
when the shared collector can own snapshot capture inside the same
budget without delaying rollback.

Advisor architecture and follow-up rationale:

- `docker-gpu-patch.ts` grows 58 lines to keep token validation before
container mutation and bounded failed-clone capture before rollback.
Splitting this security-critical ordering during the release-blocker fix
would add cross-module state transfer; extract it when the legacy swap
is retired after WSL and Jetson native proof.
- `docker-gpu-local-inference.test.ts` grows 32 lines so bridge-probe
routing assertions stay beside the behavior under test. Extract a
bridge-probe module and focused test file if that surface grows again.
- `docker-gpu-local-inference.ts` grows 15 lines to keep the
bridge-probe/host-network decision beside its caller-facing contract;
extract it with the tests if that surface grows again.
- `docker-gpu-patch.test.ts` grows 13 lines and remains below 1,350
lines; split the mode-routing cases on the next growth.
- The live fixture now supplies the canonical comma-delimited
placeholder transport. Whitespace compatibility remains covered by
parser unit tests and the messaging-provider scenario; the live proof
intentionally matches the exact canonical startup environ token.
- Dedicated `OPENSHELL_TLS_KEY` tests prove exact runtime acceptance,
arbitrary/PEM/relative/near-miss rejection, persisted `.env` rejection,
and continued rejection of supervisor identity tokens. The exact allowed
path is sourced from `NVIDIA/OpenShell@v0.0.71`
(`a242f84bb367d6df7d4d133e95a93857406c67f7`), where
`driver_utils.rs::TLS_KEY_MOUNT_PATH` defines
`/etc/openshell/tls/client/tls.key` and the Docker/Podman drivers inject
it.

This PR does not claim #6110 resolved until the reporter-class DGX Spark
aarch64 or DGX Station GB300 NVIDIA Endpoints path passes. No such
runner is declared in this repository, and organization runner inventory
is not visible with the current permissions. Missing reporter hardware
is an external acceptance blocker, not a passing result.

The #6155 docs regression fix and current `main` through `9fe45362` are
integrated. A refreshed pairwise merge-tree audit at `d76f1647` is clean
with #5595 and #6153. #5876 directly conflicts in `e2e.yaml` and related
Hermes/docs/workflow-boundary files; its resolution must union
`hermes-gpu-startup` and `mcp-bridge-dev` selectors/result summaries and
recompute uploader validation. #6020 already conflicts with current
`main` and also overlaps #6142 outside `e2e.yaml`; #6053 is mergeable
with `main` but conflicts pairwise in the uploader boundary. Those later
branches must preserve #6142's explicit-only inventory and artifact
contracts during retargeting; #6142 itself remains mergeable/CLEAN.

---
Signed-off-by: Aaron Erickson <aerickson@nvidia.com>


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **New Features**
* Strengthened Hermes GPU startup proof and managed startup integrity
assertions, plus added a skipped-by-default GPU live E2E run for startup
readiness.
* **Bug Fixes**
* Tightened PID 1 identity validation to block runtime mutation under a
foreign PID 1.
* Improved Docker GPU/OpenShell sandbox command and placeholder
handling; enhanced GPU failure diagnostics with safer redaction.
* **Documentation**
* Refined GPU passthrough and `NEMOCLAW_DOCKER_GPU_PATCH` guidance
across native Linux, Docker Desktop WSL, and Jetson/Tegra.
* **Tests**
* Expanded unit/E2E coverage for placeholder parsing, readiness refusal,
env/secret boundary enforcement, and diagnostic redaction.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Aaron Erickson <aerickson@nvidia.com>
Co-authored-by: Prekshi Vyas <prekshiv@nvidia.com>
Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
Co-authored-by: Carlos Villela <cvillela@nvidia.com>
@sandl99 sandl99 closed this Jul 3, 2026
@sandl99
sandl99 deleted the test/e2e-messaging-refactor-6022 branch July 8, 2026 03:19
Hadar301 pushed a commit to Hadar301/NemoClaw-OpenShift that referenced this pull request Jul 12, 2026
## Summary

Routes ordinary native CDI Linux through OpenShell 0.0.71's native
`--gpu` path instead of the legacy Docker container swap. The legacy
path remains available for WSL, Jetson, and explicit
`NEMOCLAW_DOCKER_GPU_PATCH=1`, with its OpenShell supervisor command
boundary and rollback diagnostics hardened.

## Related Issue

Related to NVIDIA#6110

## Changes

- Use native OpenShell GPU injection by default on ordinary native CDI
Linux.
- Document the `NEMOCLAW_DOCKER_GPU_PATCH` auto, forced-legacy, and
native-routing behavior for ordinary native Linux, Docker Desktop WSL,
and Jetson/Tegra.
- Keep `NEMOCLAW_DOCKER_GPU_PATCH=1` as the explicit legacy-swap force
control and `=0` as the existing native opt-out; Docker Desktop WSL
still ignores `=0`, and Jetson keeps its compatibility default.
- Preserve OpenShell's supervisor entrypoint on the legacy swap: Docker
receives no command tail, while the workload stays in
`OPENSHELL_SANDBOX_COMMAND`.
- Validate legacy startup tokens before stopping or renaming the
original container and serialize extra-placeholder keys as one
comma-delimited token.
- Defer every legacy recreate through the same supervisor-wait/finalize
boundary so failed clones are captured before rollback on both create
timing paths.
- Persist only allowlisted/redacted failed-clone topology, state,
process, network, and log evidence, with a 10-second total / 2-second
per-call budget so diagnostics cannot materially delay rollback.
- Permit only the canonical OpenShell Docker/Podman TLS key path in the
Hermes runtime environment; arbitrary values and persisted `.env`
entries remain rejected.
- Refresh the Dockerfile integrity pin for the changed validator so
production Hermes images fail closed on any later digest drift.
- Prove native and legacy Docker command boundaries separately,
including Ready/CUDA status, supervisor PID 1, placeholder transport,
config hashes, no backup-container leak, and inference requests.

## Type of Change

- [ ] Code change (feature, bug fix, or refactor)
- [x] Code change with doc updates
- [ ] Doc only (prose changes, no code sample modifications)
- [ ] Doc only (includes code sample changes)

## Quality Gates

- [x] Tests added or updated for changed behavior
- [ ] Existing tests cover changed behavior — justification:
- [ ] Tests not applicable — justification:
- [x] Docs updated for user-facing behavior changes
- [ ] Docs not applicable — justification:
- [x] Sensitive paths changed (security, policy, credentials, preflight,
onboarding, inference, runner, sandbox, or messaging)
- [x] Sensitive-path review completed or maintainer-approved waiver
recorded — reviewer/approval link/justification: independent exact-diff
review found no remaining substantive issue after startup-envelope
secret redaction, bounded pre-rollback capture, unified finalize
ordering, and route-specific live assertions
- [ ] Non-success, skipped, or missing CI check accepted by maintainer —
check name, approval link, and follow-up issue:

## Verification

- [x] PR description includes the DCO sign-off declaration and every
commit appears as `Verified` in GitHub
- [ ] Git hooks passed during commit and push, or `npx prek run
--from-ref main --to-ref HEAD` passes
- [x] Targeted tests pass for changed behavior
- [ ] Full `npm test` passes (broad runtime changes only)
- [x] Quality Gates section completed with required justifications or
waivers
- [x] No secrets, API keys, or credentials committed
- [ ] `npm run docs` builds without warnings (doc changes only)
- [x] Doc pages follow the [style
guide](https://github.com/NVIDIA/NemoClaw/blob/main/docs/CONTRIBUTING.md)
(doc changes only)
- [ ] New doc pages include SPDX header and frontmatter (new pages only)

Local exact-candidate verification at
`d76f1647a5f354ba04737a6a049b82bfbf6d5454`:

- CLI build and typecheck passed.
- Hermes GPU support/client/workflow coverage passed 48/48; the built
gateway-cleanup module resolves through Node, and runtime cleanup,
registration removal, and bind availability remain fail-closed.
- The shared Docker GPU diagnostic collector owns redaction for every
text/JSON artifact and returned summary; direct conventional `*_KEY` and
custom-placeholder canaries, JSON-validity, inspect-before-write,
exhausted-budget, and collector-owned top regressions pass.
- All 12 Docker GPU suites pass 126/126; the exact-head focused
E2E-support set passes 39/39, including six process-token self-match
regressions, the scrubbed integrity-proof environment, and total
forbidden-marker count.
- Conditional scan, source-shape, test-size, Biome, repository checks,
commit hooks, and push typecheck passed for the final harness
correction; only the documented macOS-invalid full CLI hook lane was
excluded.
- The forbidden-marker request sensor support suite passed 14/14 and
records counts only, never raw request bodies or marker values.
- OpenShell transport-boundary coverage passed 4/4; Docker GPU
command-envelope coverage passed 6/6; extra-placeholder parsing coverage
passed 16/16.
- Hermes validator and wrapper integrity pins match their source SHA256
digests; hadolint and diff checks pass.
- Hermes startup/boundary coverage passed 43/43 locally; the Linux-only
wrapper cases are delegated to exact-head CI.
- The full local CLI hook is not a valid gate on this macOS Node 22
host: unchanged `a1fc52c7` TypeScript entrypoints fail through the
CommonJS preload with `ERR_UNKNOWN_FILE_EXTENSION`; exact-head Linux CI
remains required.
- Hermes runtime-guard plus current-main docs regression tests: 24/24
passed.
- Hermes workflow-boundary test passed.
- Project-boundary, project-membership, test-title, source-shape,
test-size, targeted Biome, and diff checks passed.
- `npm run docs:sync-agent-variants`, `npm run
docs:check-agent-variants`, and `npm run docs` pass; Fern reports 0
errors and 2 existing warnings.

Live A/B evidence:

- Forced legacy swap at `97e3e7e1`: [run
28554699811](https://github.com/NVIDIA/NemoClaw/actions/runs/28554699811)
reproduced sustained OpenShell `Error` followed by safe rollback. It
also exposed and now closes a lifecycle instrumentation gap: the
post-create `ensureApplied()` branch bypassed pre-rollback capture.
- Native OpenShell route at signed diagnostic SHA `15f50182`: [run
28555110558](https://github.com/NVIDIA/NemoClaw/actions/runs/28555110558)
completed onboarding with exit 0, reached `Phase: Ready`, reported `CUDA
verified`, and sent authenticated Hermes chat-completions requests to
the hermetic inference endpoint. The job's only failure was the
now-fixed test regex not stripping ANSI around `Ready`.
- Prior native evidence at `7cb219d9`: [full run
28559814959](https://github.com/NVIDIA/NemoClaw/actions/runs/28559814959)
and [second pass
28559816026](https://github.com/NVIDIA/NemoClaw/actions/runs/28559816026)
both reached Ready/CUDA with clean runtime and teardown; their Hermes
proof stopped on the now-fixed ANSI matcher before downstream
assertions.
- Prior forced-legacy diagnostic on production parent `a1fc52c7` plus
workflow-only child `6d0cf6a5`: [run
28561607207](https://github.com/NVIDIA/NemoClaw/actions/runs/28561607207)
selected the legacy swap and rolled back cleanly, but failed because the
Hermes boundary rejected the driver-owned OpenShell `OPENSHELL_TLS_KEY`
path. Candidate `54cf259d` adds an exact runtime-only allowance with
negative boundary tests.
- Prior exact-head set at `a1fc52c7`: [run
28561548749](https://github.com/NVIDIA/NemoClaw/actions/runs/28561548749)
proved the GPU and security companion jobs, while Hermes GPU stopped at
a sandbox-user `/proc` permission probe after Ready/CUDA. Candidate
`54cf259d` keeps the same proof but runs it as root and restricts the
match to the exact `nemoclaw-start` process.
- Prior second native pass at `a1fc52c7`: [run
28561555945](https://github.com/NVIDIA/NemoClaw/actions/runs/28561555945)
reproduced only the same harness permission failure after Ready/CUDA,
correct PID 1 topology, authenticated inference, zero forbidden-marker
matches, and clean teardown.
- Superseded six-job set at `54cf259d`: [run
28564960504](https://github.com/NVIDIA/NemoClaw/actions/runs/28564960504)
exposed the stale Dockerfile validator digest and was canceled before
runtime proof. Candidate `970803a4` updates the integrity pin, retained
by final head `c5a67c4c`.
- Superseded second native pass at `54cf259d`: [run
28564973806](https://github.com/NVIDIA/NemoClaw/actions/runs/28564973806)
was canceled during pre-cleanup and supplies no acceptance evidence.
- Superseded forced-legacy proof on production parent `54cf259d` plus
child `69f4e1b2`: [run
28564983760](https://github.com/NVIDIA/NemoClaw/actions/runs/28564983760)
was canceled during pre-cleanup and supplies no acceptance evidence.
- Superseded six-job set at `970803a4`: [run
28565197328](https://github.com/NVIDIA/NemoClaw/actions/runs/28565197328)
was canceled before acceptance execution when the canonical
placeholder-format advisor fix advanced the head.
- Superseded second native pass at `970803a4`: [run
28565207911](https://github.com/NVIDIA/NemoClaw/actions/runs/28565207911)
was canceled before runner assignment and supplies no acceptance
evidence.
- Superseded forced-legacy proof on production parent `970803a4` plus
child `b4d5679e`: [run
28565222881](https://github.com/NVIDIA/NemoClaw/actions/runs/28565222881)
was canceled before runner assignment and supplies no acceptance
evidence.
- Superseded six-job set at `7335903b`: [run
28565576066](https://github.com/NVIDIA/NemoClaw/actions/runs/28565576066)
was intentionally canceled when the documentation gap advanced the
candidate; the root-entrypoint smoke passed, but the remaining lanes
provide no complete acceptance proof.
- Superseded second native pass at `7335903b`: [run
28565587094](https://github.com/NVIDIA/NemoClaw/actions/runs/28565587094)
was canceled before acceptance execution and supplies no acceptance
evidence.
- Superseded forced-legacy proof on production parent `7335903b` plus
child `091e16fd`: [run
28565603460](https://github.com/NVIDIA/NemoClaw/actions/runs/28565603460)
was canceled before acceptance execution and supplies no acceptance
evidence.
- Superseded six-job set at `c5a67c4c`: [run
28566083673](https://github.com/NVIDIA/NemoClaw/actions/runs/28566083673)
reached the native GPU/runtime proofs before the obsolete raw
strict-hash assertion failed; messaging independently hit the
process-probe self-match fixed by merged NVIDIA#6167. GPU, root-entrypoint,
secret-boundary, and credential companion lanes passed.
- Superseded second native pass at `c5a67c4c`: [run
28566083641](https://github.com/NVIDIA/NemoClaw/actions/runs/28566083641)
proved native routing, Ready/CUDA, `nvidia-smi`, `/proc`, `cuInit(0)=0`,
PID 1, authenticated inference, and cleanup, then failed only the
obsolete raw strict-hash assertion.
- Superseded forced-legacy proof on production parent `c5a67c4c` plus
child `48c46a7f`: [run
28566083589](https://github.com/NVIDIA/NemoClaw/actions/runs/28566083589)
proved the same runtime boundary on the legacy route, then failed only
the obsolete raw strict-hash assertion.
- Superseded six-job set at `a04a70ac`: [run
28568069499](https://github.com/NVIDIA/NemoClaw/actions/runs/28568069499)
exposed a pre-onboarding harness defect: direct Vitest import of the
production cleanup helper could not resolve its lazy CommonJS TypeScript
dependencies. Companion results do not count as final-head evidence.
- Superseded second native pass at `a04a70ac`: [run
28568069558](https://github.com/NVIDIA/NemoClaw/actions/runs/28568069558)
failed at the same pre-onboarding cleanup boundary and supplies no
runtime acceptance evidence.
- Superseded forced-legacy proof on production parent `a04a70ac` plus
child `085a3b7d`: [run
28568069530](https://github.com/NVIDIA/NemoClaw/actions/runs/28568069530)
failed at the same pre-onboarding cleanup boundary and supplies no
runtime acceptance evidence.
- Superseded six-job set at `6ac4ebc8`: [run
28568490864](https://github.com/NVIDIA/NemoClaw/actions/runs/28568490864)
exposed a clean-runner preinstall edge: the compiled cleanup child was
invoked before OpenShell existed and failed before onboarding. Companion
results do not count as final-head evidence.
- Superseded second native pass at `6ac4ebc8`: [run
28568494928](https://github.com/NVIDIA/NemoClaw/actions/runs/28568494928)
failed at the same pre-onboarding cleanup boundary and supplies no
runtime acceptance evidence.
- Superseded forced-legacy proof on production parent `6ac4ebc8` plus
child `5d3742a7`: [run
28568501430](https://github.com/NVIDIA/NemoClaw/actions/runs/28568501430)
failed at the same pre-onboarding cleanup boundary and supplies no
runtime acceptance evidence.
- Superseded six-job set at `65b06d64`: [run
28568954862](https://github.com/NVIDIA/NemoClaw/actions/runs/28568954862)
passed all six jobs, but the candidate advanced to close the
advisor-confirmed shared diagnostic-redaction boundary and two
proof-hardening review threads.
- Superseded second native pass at `65b06d64`: [run
28568959028](https://github.com/NVIDIA/NemoClaw/actions/runs/28568959028)
passed the full native runtime proof but is not final-head evidence.
- Superseded forced-legacy proof on production parent `65b06d64` plus
child `2c6dca1b`: [run
28568966557](https://github.com/NVIDIA/NemoClaw/actions/runs/28568966557)
passed but is not final-parent evidence.
- Final six-job exact-head set at `d76f1647`: [run
28601346031](https://github.com/NVIDIA/NemoClaw/actions/runs/28601346031)
passed all six requested jobs. Native GPU, Hermes startup,
root-entrypoint, secret-boundary, credential-sanitization, and messaging
proofs are green; all 21 messaging raw-token surface probes are
`ABSENT`, and every cleanup record has zero failures.
- Final second native Hermes GPU pass at `d76f1647`: [run
28601348023](https://github.com/NVIDIA/NemoClaw/actions/runs/28601348023)
passed 9/9 assertions. Artifact `8043702467`
(`sha256:0ef929fa2478f5c9579ad2f282c46de8fe4d6eefb04acb062de1af29b4eb002c`)
proves native routing, Ready/CUDA, `nvidia-smi`, `/proc`, successful
`cuInit(0)`, OpenShell PID 1, one container/no backup, authenticated
inference with zero forbidden-marker matches, and clean teardown.
- Final failed-clone rollback proof checks out exact production SHA
`d76f1647` from signed workflow-only child `4c49b5bc`: [run
28602166456](https://github.com/NVIDIA/NemoClaw/actions/runs/28602166456)
passed. Artifact `8044114375`
(`sha256:b5cff9cc7e7cf361f55e074a265c8cfefd3e6dc21ad0d300b140d2a749cde00b`)
records clone exit 137 with `failure_kind=patched_container_failed` and
`rolled_back=no` before finalize, then `rolled_back=yes`, exactly one
running original container, no backup leak, guard-observed clone
removal, clean canary scans, and clean fixture teardown.
- Final forced-legacy success proof on exact production parent
`d76f1647` plus signed workflow-only child `078a372d`: [run
28603335692](https://github.com/NVIDIA/NemoClaw/actions/runs/28603335692)
passed 9/9 assertions. Artifact `8044550700`
(`sha256:56d97c5caa8536ebff735ff600399ac7fafa2fed71206de1f5470e7d15549f6f`)
proves `gpuRoute=legacy-patch`, `--device nvidia.com/gpu=all`,
Ready/CUDA with all three GPU probes, correct OpenShell PID 1/command
envelope, one container/no backup, integrity and negative guard checks,
two authenticated inference requests with zero forbidden-marker matches,
a clean artifact canary scan, and clean teardown.

Source-of-truth review for the retained compatibility path:

- **Invalid state:** the legacy swap temporarily leaves a stopped backup
and running clone with the same OpenShell sandbox ID.
- **Source boundary:** OpenShell's Docker driver reconciles container
summaries into a map keyed only by sandbox ID; 0.0.71 can let the
stopped backup overwrite the running clone and drive the gateway into
terminal `Error`.
- **Source-fix constraint:** the NemoClaw-supported OpenShell release
does not contain deterministic active-container selection. The focused
source fix is open as
[NVIDIA/OpenShell#2116](NVIDIA/OpenShell#2116)
and passes 96/96 Docker-driver tests, strict clippy, and formatting, but
is not yet released or pinned here.
- **Regression coverage:** routing tests pin native auto / forced legacy
/ WSL / Jetson behavior; recreate tests pin capture-before-finalize on
both create timing paths; the secret canary uses the actual single
`OPENSHELL_SANDBOX_COMMAND=env ...` envelope; the live test pins native
and legacy runtime topology separately.
- **Removal condition:** remove the legacy swap and its
rollback/diagnostic modules after WSL and Jetson are proven on native
OpenShell GPU injection and the supported OpenShell floor contains
deterministic duplicate-container reconciliation.
- **WSL boundary:** Docker Desktop WSL does not expose a usable native
CDI route to this flow, so WSL retains the compatibility path and
ignores `NEMOCLAW_DOCKER_GPU_PATCH=0`; routing tests lock that behavior.
Remove it when Docker Desktop exposes usable `nvidia.com/gpu` CDI
devices to the WSL distro.
- **Jetson boundary:** Tegra `/dev/nvmap` and `/dev/nvhost-*` device
ownership requires host group propagation for the non-root sandbox user;
group-add tests lock that behavior. Remove it only when the
platform/runtime supplies equivalent access without the compatibility
recreate.

Diagnostic redaction boundary:

- **Source-of-truth invariant:** `collectDockerGpuPatchDiagnostics()`
constructs the trusted per-bundle redactor, discovers conventional and
custom-placeholder values from every known/discovered full inspect
before writing, recursively redacts JSON values, and publishes every
summary, Docker, OpenShell, and pre-rollback top artifact through that
boundary.
- **Bounded pre-rollback path:** the caller contributes only additive
values discovered from the failed clone before snapshot capture so the
shared 10-second budget cannot hide opaque values; the collector still
performs its own discovery and owns every write. Direct raw-caller and
exhausted-budget regressions scan all artifacts and returned summaries.
- **Removal condition:** remove the additive pre-rollback discovery only
when the shared collector can own snapshot capture inside the same
budget without delaying rollback.

Advisor architecture and follow-up rationale:

- `docker-gpu-patch.ts` grows 58 lines to keep token validation before
container mutation and bounded failed-clone capture before rollback.
Splitting this security-critical ordering during the release-blocker fix
would add cross-module state transfer; extract it when the legacy swap
is retired after WSL and Jetson native proof.
- `docker-gpu-local-inference.test.ts` grows 32 lines so bridge-probe
routing assertions stay beside the behavior under test. Extract a
bridge-probe module and focused test file if that surface grows again.
- `docker-gpu-local-inference.ts` grows 15 lines to keep the
bridge-probe/host-network decision beside its caller-facing contract;
extract it with the tests if that surface grows again.
- `docker-gpu-patch.test.ts` grows 13 lines and remains below 1,350
lines; split the mode-routing cases on the next growth.
- The live fixture now supplies the canonical comma-delimited
placeholder transport. Whitespace compatibility remains covered by
parser unit tests and the messaging-provider scenario; the live proof
intentionally matches the exact canonical startup environ token.
- Dedicated `OPENSHELL_TLS_KEY` tests prove exact runtime acceptance,
arbitrary/PEM/relative/near-miss rejection, persisted `.env` rejection,
and continued rejection of supervisor identity tokens. The exact allowed
path is sourced from `NVIDIA/OpenShell@v0.0.71`
(`a242f84bb367d6df7d4d133e95a93857406c67f7`), where
`driver_utils.rs::TLS_KEY_MOUNT_PATH` defines
`/etc/openshell/tls/client/tls.key` and the Docker/Podman drivers inject
it.

This PR does not claim NVIDIA#6110 resolved until the reporter-class DGX Spark
aarch64 or DGX Station GB300 NVIDIA Endpoints path passes. No such
runner is declared in this repository, and organization runner inventory
is not visible with the current permissions. Missing reporter hardware
is an external acceptance blocker, not a passing result.

The NVIDIA#6155 docs regression fix and current `main` through `9fe45362` are
integrated. A refreshed pairwise merge-tree audit at `d76f1647` is clean
with NVIDIA#5595 and NVIDIA#6153. NVIDIA#5876 directly conflicts in `e2e.yaml` and related
Hermes/docs/workflow-boundary files; its resolution must union
`hermes-gpu-startup` and `mcp-bridge-dev` selectors/result summaries and
recompute uploader validation. NVIDIA#6020 already conflicts with current
`main` and also overlaps NVIDIA#6142 outside `e2e.yaml`; NVIDIA#6053 is mergeable
with `main` but conflicts pairwise in the uploader boundary. Those later
branches must preserve NVIDIA#6142's explicit-only inventory and artifact
contracts during retargeting; NVIDIA#6142 itself remains mergeable/CLEAN.

---
Signed-off-by: Aaron Erickson <aerickson@nvidia.com>


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **New Features**
* Strengthened Hermes GPU startup proof and managed startup integrity
assertions, plus added a skipped-by-default GPU live E2E run for startup
readiness.
* **Bug Fixes**
* Tightened PID 1 identity validation to block runtime mutation under a
foreign PID 1.
* Improved Docker GPU/OpenShell sandbox command and placeholder
handling; enhanced GPU failure diagnostics with safer redaction.
* **Documentation**
* Refined GPU passthrough and `NEMOCLAW_DOCKER_GPU_PATCH` guidance
across native Linux, Docker Desktop WSL, and Jetson/Tegra.
* **Tests**
* Expanded unit/E2E coverage for placeholder parsing, readiness refusal,
env/secret boundary enforcement, and diagnostic redaction.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Aaron Erickson <aerickson@nvidia.com>
Co-authored-by: Prekshi Vyas <prekshiv@nvidia.com>
Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
Co-authored-by: Carlos Villela <cvillela@nvidia.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

area: e2e End-to-end tests, nightly failures, or validation infrastructure area: messaging Messaging channels, bridges, manifests, or channel lifecycle

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Refactor messaging E2E tests by channel feature

2 participants