Skip to content

fix(e2e): stabilize hosted inference and messaging rebuild - #5760

Merged
ericksoa merged 2 commits into
mainfrom
fix/e2e-nightly-flakes-and-messaging-regression
Jun 25, 2026
Merged

ericksoa merged 2 commits into
mainfrom
fix/e2e-nightly-flakes-and-messaging-regression

Conversation

@cv

@cv cv commented Jun 25, 2026 •

Copy link
Copy Markdown
Collaborator

Summary

Stabilizes the failing nightly hosted-inference E2E checks by making live-answer assertions tolerant of harmless model-inserted whitespace while keeping prompt-echo guards. Fixes the channels remove + rebuild regression by preserving explicit empty messaging plans through rebuild staging so token-backed channels are not rediscovered from the environment.

Related Issue

Fixes #5759

Changes

  • Preserve empty post-remove messaging plans in MessagingWorkflowPlanner.buildRebuildPlanFromSandboxEntry() and stage them into NEMOCLAW_MESSAGING_PLAN_B64 during rebuild.
  • Add regression coverage for final-channel removal rebuild plans and rebuild staging of explicit empty messaging plans.
  • Add shared E2E answer assertion helpers that normalize whitespace for deterministic numeric/token replies.
  • Update legacy bash E2Es and Vitest live replacements to use whitespace-tolerant 42 and reply-token checks, including full-e2e, sandbox-operations, launchable-smoke, agent-turn-latency, issue-4462, GPU double-onboard, and OpenClaw TUI chat-correlation paths.

Type of Change

  • Code change (feature, bug fix, or refactor)
  • Code change with doc updates
  • Doc only (prose changes, no code sample modifications)
  • Doc only (includes code sample changes)

Verification

  • PR description includes the DCO sign-off declaration and every commit appears as Verified in GitHub
  • Git hooks passed during commit and push, or npx prek run --from-ref main --to-ref HEAD passes
  • Targeted tests pass for changed behavior
  • Full npm test passes (broad runtime changes only)
  • Tests added or updated for new or changed behavior
  • No secrets, API keys, or credentials committed
  • Docs updated for user-facing behavior changes
  • npm run docs builds without warnings (doc changes only)
  • Doc pages follow the style guide (doc changes only)
  • New doc pages include SPDX header and frontmatter (new pages only)

Verification details:

  • npm run build:cli
  • npx vitest run --project cli test/helpers/e2e-answer-assertions.test.ts test/openclaw-tui-chat-correlation.test.ts src/lib/messaging/compiler/workflow-planner.test.ts src/lib/actions/sandbox/rebuild-messaging-stage.test.ts
  • npm run typecheck:cli
  • npx prek run shellcheck --files test/e2e/lib/openclaw-json.sh test/e2e/test-full-e2e.sh test/e2e/test-sandbox-operations.sh test/e2e/test-launchable-smoke.sh test/e2e/test-agent-turn-latency-e2e.sh test/e2e/test-issue-4462-scope-upgrade-approval.sh
  • Commit and push hooks passed, including CLI tests and CLI TypeScript pre-push checks.
  • Documentation writer review: no docs changes needed because this is E2E harness stabilization plus preserving the already-documented channel remove/rebuild behavior.

Signed-off-by: Carlos Villela cvillela@nvidia.com

Summary by CodeRabbit

  • Bug Fixes

    • Improved messaging rebuild planning so an explicitly empty channel set is preserved (not cleared or re-derived).
    • Made reply-token and integer 42 validations more robust to harmless whitespace/newlines in streamed output.
    • Enhanced chat-correlation checks to better tolerate formatting variations.
  • Tests

    • Added/extended unit coverage for rebuild planning and “empty channels” behavior.
    • Introduced shared E2E assertion helpers and updated live/shell E2E scenarios to use them.

Signed-off-by: Carlos Villela <cvillela@nvidia.com>
@cv cv self-assigned this Jun 25, 2026
@coderabbitai

coderabbitai Bot commented Jun 25, 2026 •

Copy link
Copy Markdown
Contributor

Review Change Stack

No actionable comments were generated in the recent review. 🎉

ℹ️ Recent review info
⚙️ Run configuration

Configuration used: Path: .coderabbit.yaml

Review profile: CHILL

Plan: Enterprise

Run ID: 8e30b8c3-d785-47fb-8110-ff736f932d76

📥 Commits

Reviewing files that changed from the base of the PR and between 0cfd590 and 5105c33.

📒 Files selected for processing (1)
  • test/e2e-scenario/live/issue-4462-scope-upgrade-approval.test.ts
🚧 Files skipped from review as they are similar to previous changes (1)
  • test/e2e-scenario/live/issue-4462-scope-upgrade-approval.test.ts

📝 Walkthrough

Walkthrough

The PR preserves explicit empty messaging rebuild plans after channel removal and replaces exact 42/reply-token matching with whitespace-tolerant helpers across Vitest and shell E2E scenarios.

Changes

Messaging rebuild and E2E assertions

Layer / File(s) Summary
Rebuild planner preserves empty plan
src/lib/messaging/compiler/workflow-planner.ts, src/lib/messaging/compiler/workflow-planner.test.ts
The rebuild planner keeps an explicit empty channels: [] plan after allowlist filtering, and the regression test covers the empty rebuild shape after removing the last channel.
Sandbox stages empty stored plan
src/lib/actions/sandbox/rebuild-messaging-stage.test.ts
Adds an explicit empty stored messaging-plan fixture and verifies that rebuild staging writes the empty plan without clearing the env or rediscovering channels.
Shared answer helpers
test/helpers/e2e-answer-assertions.ts, test/helpers/e2e-answer-assertions.test.ts
Adds whitespace-stripping helpers for integer 42 and reply-token matching, with unit tests for normalization and detection behavior.
42 assertions use shared helper
test/e2e-scenario/live/agent-turn-latency.test.ts, test/e2e-scenario/live/full-e2e.test.ts, test/e2e-scenario/live/gpu-double-onboard.test.ts, test/e2e-scenario/live/launchable-smoke.test.ts, test/e2e-scenario/live/sandbox-operations.test.ts, test/e2e-scenario/live/issue-4462-scope-upgrade-approval.test.ts, test/e2e/lib/openclaw-json.sh, test/e2e/test-agent-turn-latency-e2e.sh, test/e2e/test-full-e2e.sh, test/e2e/test-issue-4462-scope-upgrade-approval.sh, test/e2e/test-launchable-smoke.sh, test/e2e/test-sandbox-operations.sh
The live Vitest scenarios and shell E2E scripts replace inline 42 regex checks with shared helper-based assertions, including the embedded shell logic in the issue-4462 live scenario.
Reply-token whitespace correlation
test/openclaw-tui-chat-correlation.test.ts, test/e2e-scenario/live/openclaw-tui-chat-correlation.test.ts
The OpenClaw TUI chat correlation tests match reply tokens after whitespace compaction in trace analysis, the live repro script, and the split-token regression test.

Possibly related PRs

  • NVIDIA/NemoClaw#5673 — Both changes adjust workflow-planner channel filtering behavior around empty or deny-all channel sets during rebuild-related planning.
  • NVIDIA/NemoClaw#5743 — Both touch the sandbox messaging rebuild path around stageMessagingManifestPlanForRebuild and stored plan handling.

Suggested labels

bug-fix

Suggested reviewers

  • ericksoa

Estimated code review effort

🎯 3 (Moderate) | ⏱️ ~30 minutes

Poem

A bunny nibbled whitespace free,
and 42 still hopped with glee.
Empty plans stayed tucked in place,
reply tokens found their space,
hop-hop! The tests all passed for me. 🐇

🚥 Pre-merge checks | ✅ 4 | ❌ 1

❌ Failed checks (1 warning)

Check name Status Explanation Resolution
Docstring Coverage ⚠️ Warning Docstring coverage is 0.00% which is insufficient. The required threshold is 80.00%. Write docstrings for the functions missing them to satisfy the coverage threshold.
✅ Passed checks (4 passed)
Check name Status Explanation
Description Check ✅ Passed Check skipped - CodeRabbit’s high-level summary is enabled.
Title check ✅ Passed The title is concise and accurately reflects the main changes to E2E stabilization and messaging rebuild handling.
Linked Issues check ✅ Passed The changes address the linked issue's code requirements: whitespace-tolerant E2E assertions, semantic token matching, and preserving empty messaging rebuild plans.
Out of Scope Changes check ✅ Passed The modified files stay within the stated goals and do not introduce clearly unrelated code changes.

✏️ Tip: You can configure your own custom pre-merge checks in the settings.

✨ Finishing Touches
📝 Generate docstrings
  • Create stacked PR
  • Commit on current branch
🧪 Generate unit tests (beta)
  • Create PR with unit tests
  • Commit unit tests in branch fix/e2e-nightly-flakes-and-messaging-regression

Comment @coderabbitai help to get the list of available commands.

@github-code-quality

github-code-quality Bot commented Jun 25, 2026 •

Copy link
Copy Markdown
Contributor

Code Coverage Overview

Languages: TypeScript

TypeScript / code-coverage/plugin

The overall coverage in the fix/e2e-nightly-flak... branch is 96%. Coverage data for the main branch is not yet available.

Show a code coverage summary of the most covered files.
File main fix/e2e-nightly-flak... 5105c33 +/-
nemoclaw/src/se...cret-scanner.ts — 100% —
nemoclaw/src/commands/slash.ts — 100% —
nemoclaw/src/li...bprocess-env.ts — 100% —
nemoclaw/src/bl...eprint/state.ts — 98% —
nemoclaw/src/onboard/config.ts — 98% —
nemoclaw/src/bl...int/snapshot.ts — 97% —
nemoclaw/src/bl...print/runner.ts — 95% —
nemoclaw/src/co...ration-state.ts — 94% —
nemoclaw/src/bl...ate-networks.ts — 94% —
nemoclaw/src/index.ts — 94% —

TypeScript / code-coverage/cli

The overall coverage in the fix/e2e-nightly-flak... branch is 47%. Coverage data for the main branch is not yet available.

Show a code coverage summary of the most covered files.
File main fix/e2e-nightly-flak... 5105c33 +/-
src/lib/state/o...oard-session.ts — 91% —
src/lib/inference/local.ts — 76% —
src/lib/sandbox/config.ts — 72% —
src/lib/actions...dbox/rebuild.ts — 71% —
src/lib/onboard/preflight.ts — 64% —
src/lib/actions...licy-channel.ts — 60% —
src/lib/state/sandbox.ts — 55% —
src/lib/onboard...er-gpu-patch.ts — 50% —
src/lib/policy/index.ts — 49% —
src/lib/onboard.ts — 19% —

Updated June 25, 2026 03:22 UTC
Code Coverage is in Public Preview. Learn more and provide us with your feedback.

@github-actions

github-actions Bot commented Jun 25, 2026 •

Copy link
Copy Markdown
Contributor

PR Review Advisor — No blocking findings

Merge posture: No blocking advisor findings
Primary next action: Add or justify PRA-T1 and any related test follow-ups.
Open items: 0 required · 0 warnings · 0 suggestions · 8 test follow-ups
Since last review: 0 prior items resolved · 7 still apply · 0 new items found

Action checklist

  • PRA-T1 Add or justify test follow-up: Runtime validation
  • PRA-T2 Add or justify test follow-up: Runtime validation
  • PRA-T3 Add or justify test follow-up: Runtime validation
  • PRA-T4 Add or justify test follow-up: Runtime validation
  • PRA-T5 Add or justify test follow-up: Runtime validation
  • PRA-T6 Add or justify test follow-up: Acceptance clause
  • PRA-T7 Add or justify test follow-up: Acceptance clause
  • PRA-T8 Add or justify test follow-up: Acceptance clause
Test follow-ups to resolve or justify

If these cover changed behavior, prefer adding them in this PR; otherwise state why existing coverage is enough or link the follow-up.

  • PRA-T1 Runtime validation — Run or add an automated channels lifecycle scenario where TELEGRAM_BOT_TOKEN remains exported, `channels remove telegram` succeeds, then `rebuild` leaves registry `messaging.plan.channels`, OpenClaw config, gateway providers, and policy presets without `telegram`.. Unit and staging coverage are focused and adequate for the code change, but the highest-confidence proof for the original regression is a real channel lifecycle through sandbox rebuild with the original token environment still present.
  • PRA-T2 Runtime validation — Run or identify the `channels-add-remove-e2e` evidence for the remove-plus-rebuild path; do not report external job pass/fail in this code-review surface.. Unit and staging coverage are focused and adequate for the code change, but the highest-confidence proof for the original regression is a real channel lifecycle through sandbox rebuild with the original token environment still present.
  • PRA-T3 Runtime validation — Run or identify the `cloud-e2e`/full-E2E evidence that a parsed OpenClaw answer of `4\n2` is accepted while prompt/fallback/transport-error guards still fail when present.. Unit and staging coverage are focused and adequate for the code change, but the highest-confidence proof for the original regression is a real channel lifecycle through sandbox rebuild with the original token environment still present.
  • PRA-T4 Runtime validation — Run or identify the `launchable-smoke-e2e` evidence that the launchable OpenClaw agent path accepts whitespace-split `42` after successful inference routing.. Unit and staging coverage are focused and adequate for the code change, but the highest-confidence proof for the original regression is a real channel lifecycle through sandbox rebuild with the original token environment still present.
  • PRA-T5 Runtime validation — Run or identify the `sandbox-operations-e2e` evidence that TC-SBX-02 accepts whitespace-split `42` only from parsed agent reply text.. Unit and staging coverage are focused and adequate for the code change, but the highest-confidence proof for the original regression is a real channel lifecycle through sandbox rebuild with the original token environment still present.
  • PRA-T6 Acceptance clause — rebuild must not rediscover `TELEGRAM_BOT_TOKEN` after `channels remove telegram`. — add test evidence or identify existing coverage. Static/unit evidence covers the core cause: final-channel removal produces an explicit empty rebuild plan with no channels, credentials, or policy entries, and staging writes it instead of clearing the env. Full runtime proof still needs a channel lifecycle run where TELEGRAM_BOT_TOKEN remains exported through remove plus rebuild.
  • PRA-T7 Acceptance clause — Re-run at least: — add test evidence or identify existing coverage. External E2E execution status is outside this code-review surface. The diff contains targeted unit/staging coverage and updates the named E2E assertion sites.
  • PRA-T8 Acceptance clause — `channels-add-remove-e2e` — add test evidence or identify existing coverage. External E2E execution status is not evaluated here; planner and rebuild staging tests cover the channel remove plus rebuild plan transformation that underlies this scenario.

Workflow run details

This is an automated, non-binding review; it still expects maintainers and agents to respond to each required or warning item. Treat suggestions as current-PR improvements when they touch changed code; defer only with maintainer rationale or a linked follow-up. A human maintainer must make the final merge decision.

@github-actions

github-actions Bot commented Jun 25, 2026 •

Copy link
Copy Markdown
Contributor

E2E Advisor Recommendation

Required E2E: channels-add-remove-vitest
Optional E2E: sandbox-rebuild-vitest, channels-stop-start-vitest, messaging-providers-vitest

Dispatch hint: channels-add-remove-vitest

Workflow run

Full advisor summary

E2E Recommendation Advisor

Base: origin/main
Head: HEAD
Confidence: high

Required E2E

  • channels-add-remove-vitest (high; real Docker/OpenShell sandbox with hosted inference credential and rebuilds, timeout 75 minutes): Required because the runtime planner change affects the exact real-user flow where OpenClaw is onboarded without messaging, Telegram is added, the sandbox is rebuilt, Telegram is removed, and the sandbox is rebuilt again. This job verifies the final empty messaging plan does not rediscover token-backed channels and that provider, network policy, registry plan, and rendered openclaw.json state are clean after rebuild.

Optional E2E

  • sandbox-rebuild-vitest (high; real Docker/OpenShell sandbox rebuild, timeout 90 minutes): Useful adjacent confidence for the general real sandbox rebuild lifecycle, although it is less targeted than channels-add-remove for the empty messaging-plan regression.
  • channels-stop-start-vitest (high; real OpenClaw/Hermes channel stop/start and rebuild matrix, timeout 90 minutes): Optional coverage for rebuilds of existing messaging plans with disabled/enabled channels and provider reuse. It is adjacent to the planner code but does not specifically cover removing the final channel.
  • messaging-providers-vitest (high; live messaging provider/credential-policy chain, timeout 90 minutes): Optional confidence for token-backed messaging credential provider, placeholder, L7 proxy, and policy behavior after the shared messaging planner changes, but the changed behavior is most directly covered by channels-add-remove.

New E2E recommendations

  • None.

Dispatch hint

  • Workflow: E2E / Vitest Scenarios
  • jobs input: channels-add-remove-vitest

@github-actions

github-actions Bot commented Jun 25, 2026 •

Copy link
Copy Markdown
Contributor

Vitest E2E Scenario Recommendation

Required Vitest E2E scenarios: agent-turn-latency-vitest, full-e2e-vitest, gpu-double-onboard-vitest, issue-4462-scope-upgrade-approval-vitest, launchable-smoke-vitest, openclaw-tui-chat-correlation-vitest
Optional Vitest E2E scenarios: None

Dispatch required Vitest E2E scenarios:

  • gh workflow run e2e-vitest-scenarios.yaml --ref <pr-head-ref> --field jobs=agent-turn-latency-vitest
  • gh workflow run e2e-vitest-scenarios.yaml --ref <pr-head-ref> --field jobs=full-e2e-vitest
  • gh workflow run e2e-vitest-scenarios.yaml --ref <pr-head-ref> --field jobs=gpu-double-onboard-vitest
  • gh workflow run e2e-vitest-scenarios.yaml --ref <pr-head-ref> --field jobs=issue-4462-scope-upgrade-approval-vitest
  • gh workflow run e2e-vitest-scenarios.yaml --ref <pr-head-ref> --field jobs=launchable-smoke-vitest
  • gh workflow run e2e-vitest-scenarios.yaml --ref <pr-head-ref> --field jobs=openclaw-tui-chat-correlation-vitest

Workflow run

Full Vitest E2E advisor summary

Vitest E2E Scenario Advisor

Base: origin/main
Head: HEAD
Confidence: high

Required Vitest E2E scenarios

  • agent-turn-latency-vitest: Focused free-standing Vitest job wired for changed live test test/e2e-scenario/live/agent-turn-latency.test.ts.
    • Dispatch: gh workflow run e2e-vitest-scenarios.yaml --ref <pr-head-ref> --field jobs=agent-turn-latency-vitest
  • full-e2e-vitest: Focused free-standing Vitest job wired for changed live test test/e2e-scenario/live/full-e2e.test.ts.
    • Dispatch: gh workflow run e2e-vitest-scenarios.yaml --ref <pr-head-ref> --field jobs=full-e2e-vitest
  • gpu-double-onboard-vitest: Focused free-standing Vitest job wired for changed live test test/e2e-scenario/live/gpu-double-onboard.test.ts.
    • Dispatch: gh workflow run e2e-vitest-scenarios.yaml --ref <pr-head-ref> --field jobs=gpu-double-onboard-vitest
  • issue-4462-scope-upgrade-approval-vitest: Focused free-standing Vitest job wired for changed live test test/e2e-scenario/live/issue-4462-scope-upgrade-approval.test.ts.
    • Dispatch: gh workflow run e2e-vitest-scenarios.yaml --ref <pr-head-ref> --field jobs=issue-4462-scope-upgrade-approval-vitest
  • launchable-smoke-vitest: Focused free-standing Vitest job wired for changed live test test/e2e-scenario/live/launchable-smoke.test.ts.
    • Dispatch: gh workflow run e2e-vitest-scenarios.yaml --ref <pr-head-ref> --field jobs=launchable-smoke-vitest
  • openclaw-tui-chat-correlation-vitest: Focused free-standing Vitest job wired for changed live test test/e2e-scenario/live/openclaw-tui-chat-correlation.test.ts.
    • Dispatch: gh workflow run e2e-vitest-scenarios.yaml --ref <pr-head-ref> --field jobs=openclaw-tui-chat-correlation-vitest

Optional Vitest E2E scenarios

  • None.

Relevant changed files

  • src/lib/messaging/compiler/workflow-planner.ts
  • test/e2e-scenario/live/agent-turn-latency.test.ts
  • test/e2e-scenario/live/full-e2e.test.ts
  • test/e2e-scenario/live/gpu-double-onboard.test.ts
  • test/e2e-scenario/live/issue-4462-scope-upgrade-approval.test.ts
  • test/e2e-scenario/live/launchable-smoke.test.ts
  • test/e2e-scenario/live/openclaw-tui-chat-correlation.test.ts
  • test/e2e-scenario/live/sandbox-operations.test.ts
  • test/helpers/e2e-answer-assertions.ts

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 1

🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

Inline comments:
In `@test/e2e-scenario/live/issue-4462-scope-upgrade-approval.test.ts`:
- Around line 122-124: The contains_integer_42 helper is using a brittle tr |
grep pipeline that can fail under set -euo pipefail. Update contains_integer_42
to follow the captured-string approach used in test/e2e/lib/openclaw-json.sh, so
it first reads the input into a variable, strips whitespace from that value, and
then performs the integer 42 match without relying on a pipeline.
🪄 Autofix (Beta)

Fix all unresolved CodeRabbit comments on this PR:

  • Push a commit to this branch (recommended)
  • Create a new PR with the fixes

ℹ️ Review info
⚙️ Run configuration

Configuration used: Path: .coderabbit.yaml

Review profile: CHILL

Plan: Enterprise

Run ID: fb6ff966-d8a8-4ea6-a01c-845bbebb1a77

📥 Commits

Reviewing files that changed from the base of the PR and between ae94f13 and 0cfd590.

📒 Files selected for processing (19)
  • src/lib/actions/sandbox/rebuild-messaging-stage.test.ts
  • src/lib/messaging/compiler/workflow-planner.test.ts
  • src/lib/messaging/compiler/workflow-planner.ts
  • test/e2e-scenario/live/agent-turn-latency.test.ts
  • test/e2e-scenario/live/full-e2e.test.ts
  • test/e2e-scenario/live/gpu-double-onboard.test.ts
  • test/e2e-scenario/live/issue-4462-scope-upgrade-approval.test.ts
  • test/e2e-scenario/live/launchable-smoke.test.ts
  • test/e2e-scenario/live/openclaw-tui-chat-correlation.test.ts
  • test/e2e-scenario/live/sandbox-operations.test.ts
  • test/e2e/lib/openclaw-json.sh
  • test/e2e/test-agent-turn-latency-e2e.sh
  • test/e2e/test-full-e2e.sh
  • test/e2e/test-issue-4462-scope-upgrade-approval.sh
  • test/e2e/test-launchable-smoke.sh
  • test/e2e/test-sandbox-operations.sh
  • test/helpers/e2e-answer-assertions.test.ts
  • test/helpers/e2e-answer-assertions.ts
  • test/openclaw-tui-chat-correlation.test.ts

Comment thread test/e2e-scenario/live/issue-4462-scope-upgrade-approval.test.ts
@sandl99 sandl99 assigned sandl99 and unassigned sandl99 Jun 25, 2026
Signed-off-by: Carlos Villela <cvillela@nvidia.com>
@github-actions

Copy link
Copy Markdown
Contributor

Vitest E2E Scenario Results — ✅ All requested jobs passed

Run: 28144691365
Workflow ref: fix/e2e-nightly-flakes-and-messaging-regression
Requested scenarios: (default — all supported)
Requested jobs: channels-add-remove-vitest
Summary: 1 passed, 0 failed, 0 cancelled, 0 skipped

Job Result
channels-add-remove-vitest ✅ success

@github-actions

Copy link
Copy Markdown
Contributor

Vitest E2E Scenario Results — ❌ Some jobs failed

Run: 28144742533
Workflow ref: fix/e2e-nightly-flakes-and-messaging-regression
Requested scenarios: (default — all supported)
Requested jobs: agent-turn-latency-vitest,full-e2e-vitest,gpu-double-onboard-vitest,issue-4462-scope-upgrade-approval-vitest,launchable-smoke-vitest,openclaw-tui-chat-correlation-vitest
Summary: 3 passed, 3 failed, 0 cancelled, 0 skipped

Job Result
agent-turn-latency-vitest ❌ failure
full-e2e-vitest ✅ success
gpu-double-onboard-vitest ✅ success
issue-4462-scope-upgrade-approval-vitest ✅ success
launchable-smoke-vitest ❌ failure
openclaw-tui-chat-correlation-vitest ❌ failure

Failed jobs: agent-turn-latency-vitest, launchable-smoke-vitest, openclaw-tui-chat-correlation-vitest. Check run artifacts for logs.

@cv

cv commented Jun 25, 2026

Copy link
Copy Markdown
Collaborator Author

E2E follow-up for the advisor items:

  • Required channels-add-remove-vitest passed in run https://github.com/NVIDIA/NemoClaw/actions/runs/28144691365. This covers the remove+rebuild regression with token-backed channel env still present and verifies registry/config/provider/policy cleanup.
  • I also dispatched the Vitest scenario advisor set in https://github.com/NVIDIA/NemoClaw/actions/runs/28144742533:
    • Passed: full-e2e-vitest, gpu-double-onboard-vitest, issue-4462-scope-upgrade-approval-vitest.
    • Failed before reaching this PR's changed assertions: launchable-smoke-vitest and openclaw-tui-chat-correlation-vitest both failed on invalid hosted inference key setup (NVIDIA_INFERENCE_API_KEY must start with nvapi- / Invalid NVIDIA API key) during pre-onboard/onboard; agent-turn-latency-vitest failed during onboarding with a stale previous-session guard before the OpenClaw/Hermes answer assertions ran.

So the targeted runtime proof for the real messaging regression is green, and the additional live assertion lanes that failed did so before exercising the whitespace-normalized assertion changes.

@ericksoa
ericksoa merged commit e3b8325 into main Jun 25, 2026
180 of 183 checks passed
@ericksoa
ericksoa deleted the fix/e2e-nightly-flakes-and-messaging-regression branch June 25, 2026 04:48
jyaunches added a commit that referenced this pull request Jun 25, 2026
## Summary
Ports the focused E2E stabilizers from `dep/openshell-v0.0.67` / PR
#5596 onto current `main` after PR #5760, without merging the full
OpenShell 0.0.67 branch.

This targets the full-main nightly failures from run 28172043426:
- `kimi-inference-compat-e2e` — relax live Kimi trajectory shape
expectations.
- `common-egress-agent-e2e` — tolerate wrapped reply tokens like
`REFER\nENCE_AGENT_OK`.
- `sessions-agents-cli-e2e` — keep sessions admin RPCs local/SDK-backed
and avoid multiline RPC args.

Also includes the small channel/remove rebuild staging stabilizer
carried by the shared matrix-stabilization commit.

## Validation
- Local push hooks could not fully run because this worktree is missing
local npm dependencies (`tsx`, `typescript`, Biome dependency `klaw`).
- Shellcheck/gitleaks/basic pre-commit checks passed before the
dependency-gated hooks failed.
- Focused nightly E2E dispatch is being run separately on this branch.

## Notes
- Does not port the full OpenShell 0.0.67 upgrade.
- Does not claim to fix `diagnostics-e2e` HTTP 403; that failure looked
infra/upstream/credential-like.

Signed-off-by: Julie Yaunches <jyaunches@nvidia.com>


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **Bug Fixes**
* Improved sandbox gateway RPC execution with pairing-aware retry, clear
retry/no-retry gating, and stricter handling of unsupported admin
methods.
* Added safer parsing and richer failure diagnostics with token
redaction in returned output and logged errors.
* **New Features**
* Enhanced gateway RPC results to include separate diagnostic output and
tightened admin method support via allowlisting.
* **Tests**
* Expanded Vitest coverage for gateway orchestration/output handling and
stream capture behavior.
* Strengthened OpenClaw text assertions, updated e2e token/PONG checks,
and relaxed Kimi validations for mock vs live.
  * Prevented Telegram env reuse after channel removal.
* **Chores**
* Added optional stdout/stderr stream capture controls for OpenShell
helpers.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Co-authored-by: Aaron Erickson <aerickson@nvidia.com>
cv added a commit that referenced this pull request Jun 25, 2026
## Summary
Aligns the Vitest E2E lanes that receive `NVIDIA_INFERENCE_API_KEY` with
the hosted Inference Hub compatible-endpoint contract. This prevents
Build API `nvapi-` validation from rejecting the hosted-inference secret
before the affected scenarios reach their live assertions.

## Related Issue
Refs #5759

## Changes
- Configure `agent-turn-latency-vitest`, `launchable-smoke-vitest`, and
`openclaw-tui-chat-correlation-vitest` with hosted-compatible
provider/model env and `COMPATIBLE_API_KEY` sourced from
`NVIDIA_INFERENCE_API_KEY`.
- Teach the shared `cloud-openclaw` onboarding fixture to honor
`NEMOCLAW_E2E_USE_HOSTED_INFERENCE=1` by staging the compatible endpoint
provider and model.
- Update launchable smoke to skip `nvapi-` validation when running in
hosted-compatible mode and to expect the `compatible-endpoint` route.
- Add `--fresh`/`NEMOCLAW_FRESH=1` to agent-turn-latency install
attempts so retries do not trip over failed prior onboarding sessions.

## Type of Change
- [x] Code change (feature, bug fix, or refactor)
- [ ] Code change with doc updates
- [ ] Doc only (prose changes, no code sample modifications)
- [ ] Doc only (includes code sample changes)

## Verification
- [x] PR description includes the DCO sign-off declaration and every
commit appears as `Verified` in GitHub
- [x] Git hooks passed during commit and push, or `npx prek run
--from-ref main --to-ref HEAD` passes
- [x] Targeted tests pass for changed behavior
- [ ] Full `npm test` passes (broad runtime changes only)
- [x] Tests added or updated for new or changed behavior
- [x] No secrets, API keys, or credentials committed
- [ ] Docs updated for user-facing behavior changes
- [ ] `npm run docs` builds without warnings (doc changes only)
- [ ] Doc pages follow the [style
guide](https://github.com/NVIDIA/NemoClaw/blob/main/docs/CONTRIBUTING.md)
(doc changes only)
- [ ] New doc pages include SPDX header and frontmatter (new pages only)

Verification details:
- Synced with `origin/main` and confirmed the latest commit was #5760;
these Vitest lane hosted-key fixes were not already present.
- `npm run typecheck:cli`
- `npx vitest run --project cli
test/e2e-scenario/support-tests/e2e-scenarios-workflow.test.ts
test/e2e-script-workflow.test.ts
test/e2e-scenario/support-tests/hosted-inference.test.ts`
- `npx prek run check-yaml --files
.github/workflows/e2e-vitest-scenarios.yaml`
- Commit and push hooks passed, including CLI tests and CLI TypeScript
pre-push checks.
- Documentation writer review: no docs changes needed because this only
updates CI/live E2E harness behavior.

---
<!-- DCO sign-off is required in this PR description, and every commit
must appear as Verified in GitHub. Run: git config user.name && git
config user.email -->
Signed-off-by: Carlos Villela <cvillela@nvidia.com>


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **New Features**
* Added an opt-in “hosted inference compatibility” mode for end-to-end
onboarding and live scenarios, including hosted model/endpoint/provider
routing and onboarding environment setup.
* **Tests**
* Updated live smoke and helper flows to reflect hosted-compatible
behavior, including conditional API key checks and inference routing
assertions.
  * Improved test consistency with a “fresh” sandbox install option.
* **Chores**
* Updated CI end-to-end Vitest execution to run via the hosted NVIDIA
inference setup with secure API key injection and hosted model
configuration.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Carlos Villela <cvillela@nvidia.com>
jyaunches added a commit that referenced this pull request Jun 25, 2026
## Summary
Restore issue #5800 parity package `P0-B` for merged
onboard/rebuild/lifecycle bash-suite deltas only.

## Related Issues
Refs #5800
Refs #5098
Refs #5225
Refs #5487
Refs #5410
Refs #5760

## Scope gate
- Package: `P0-B — Onboard/rebuild/lifecycle parity`
- Included PRs all merged and touched `test/e2e`: yes — #5225, #5487,
#5410, #5760
- Out of scope: unmerged/non-bash PRs; shell lane retirement / PR #5756
cleanup

## Parity map
| ID | Source PR | Contract | Inference classification | Vitest
assertion / waiver | Status |
| --- | --- | --- | --- | --- | --- |
| B1 | #5225 | Persisted sandbox-entry gateway resolution is used by
lifecycle commands. | `none` | Existing
`src/lib/actions/sandbox/sandbox-gateway-routing.test.ts`,
`src/lib/onboard/gateway-binding.test.ts`,
`src/lib/onboard/sandbox-registration.test.ts` | covered |
| B2 | #5225 | Onboard repair/double-onboard failures include captured
onboard output diagnostics. | `none` | Existing live
`test/e2e-scenario/live/onboard-repair.test.ts`,
`test/e2e-scenario/live/double-onboard.test.ts`; diagnostics are
bash-runner-only verbosity and not a durable Vitest assertion. | waived
|
| B3 | #5487 | Plain `nemoclaw onboard` auto-detects an `in_progress`
session and resumes without `--resume`; `--fresh` suppresses
auto-resume. | `hosted-compatible capable` |
`test/e2e-scenario/live/onboard-resume.test.ts` Phase 3.5 mutates the
completed session to `in_progress`, asserts `(resume mode)` + cached
skips, then asserts `--fresh` fails at injected preflight without resume
banner. | covered |
| B4 | #5410 | Rebuild resumes messaging from `messaging.plan` rather
than legacy `providerCredentialHashes`; stale top-level provider hash
state is not used. | `none` |
`test/e2e-scenario/live/rebuild-hermes.test.ts` curated registry omits
`providerCredentialHashes`;
`src/lib/onboard/machine/handlers/sandbox.test.ts` refreshes
registry-plan credential hashes from env on rebuild resume. | covered |
| B5 | #5410 | Empty/staged rebuild messaging plan is preserved and
token-backed channels are not rediscovered. | `none` | Existing
`src/lib/actions/sandbox/rebuild-messaging-stage.test.ts` and
`src/lib/onboard/machine/handlers/sandbox.test.ts`. | covered |
| B6 | #5760 | Hosted-inference/messaging rebuild and live answer
assertions tolerate model whitespace around integer `42`. |
`hosted-compatible capable` | Existing
`test/helpers/e2e-answer-assertions.test.ts`, consumed by
`agent-turn-latency`, `full-e2e`, `launchable-smoke`, and
`sandbox-operations` live Vitests. | covered |
| B7 | #5760 | Stabilized hosted inference and messaging rebuild remain
validated by live full/sandbox/launchable/rebuild targets. |
`hosted-compatible capable` | Existing live targets remain unchanged;
local live execution blocked by Docker daemon unavailable. Selective
workflow required after PR opens. | follow-up |

## Inference mode support
- Default mode for touched live targets: `hosted-compatible capable` for
onboard-resume/rebuild-hermes/full/sandbox/launchable answer paths;
`none` for unit/process registry and gateway routing tests.
- Real inference support preserved: yes for existing hosted-compatible
live targets; no new inference adapter seam added here.
- Modes validated in this PR: local unit/process Vitests only; live
hosted-compatible validation needs GitHub runner/secrets because local
Docker daemon is unavailable.
- If not validated with real inference: local
`NEMOCLAW_RUN_E2E_SCENARIOS=1 npx vitest run --project
e2e-scenarios-live test/e2e-scenario/live/onboard-resume.test.ts
test/e2e-scenario/live/rebuild-hermes.test.ts` failed at prereq Docker
daemon check before scenario assertions.

## Validation
- [x] `git diff --check`
- [x] `npm test -- src/lib/onboard/machine/handlers/sandbox.test.ts
test/helpers/e2e-answer-assertions.test.ts
src/lib/actions/sandbox/sandbox-gateway-routing.test.ts
src/lib/onboard/entry-options.test.ts
src/lib/onboard/sandbox-registration.test.ts
src/lib/actions/sandbox/rebuild-messaging-stage.test.ts`
- [x] `npm run typecheck:cli`
- [x] `npm run build:cli`
- [x] `npm run test-size:check`
- [ ] hosted-compatible selective E2E workflow: pending PR / runner
dispatch

## Follow-ups / waivers
- Waiver B2: bash-only diagnostic verbosity from #5225 is not a durable
Vitest contract; existing live tests already preserve the functional
repair/double-onboard behavior.
- Follow-up B7: dispatch selective live Vitest scenarios on GitHub
runner with Docker + hosted inference secret after PR opens.


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **Tests**
* Added coverage for sandbox resume flows to verify updated Telegram
credentials are picked up when resuming a rebuild.
* Expanded end-to-end onboarding resume scenarios to confirm implicit
resume behavior, including skipped cached steps and fresh runs starting
from the expected point.
* Strengthened rebuild scenario checks to ensure curated registry
entries no longer include legacy credential-hash data.

<!-- end of auto-generated comment: release notes by coderabbit.ai -->
@cv cv added the v0.0.68 label Jun 25, 2026
jyaunches added a commit that referenced this pull request Jun 25, 2026
## Summary
Restore the Kimi-specific issue #5800 parity work for package `P0-D`;
existing recovery and scope-upgrade package rows are explicitly mapped
as pre-existing coverage and revalidated context, but not changed
acceptance scope in this PR.

## Related Issues
Refs #5800
Refs #5098
Refs #5342
Refs #5401
Refs #5406
Refs #5412
Refs #5413
Refs #5625
Refs #5760

## Scope gate
- Package: `P0-D — Recovery, Kimi, and scope-upgrade parity`
- Included PRs all merged and touched `test/e2e`: yes
- Changed acceptance scope in this PR: Kimi public-NVIDIA/mock parity
(`D2`, `D3`)
- Existing package rows revalidated without diff changes: recovery
(`D1`) and scope-upgrade (`D4`)
- Out of scope: unmerged/non-bash PRs; shell lane retirement / PR #5756
cleanup

## Parity map
| ID | Source PR | Contract | Inference classification | Vitest
assertion / waiver | Status |
| --- | --- | --- | --- | --- | --- |
| D1 | #5342, #5401 | Recovery proxy env sourcing, missing proxy-env
warning, guard retention, ciao/networkInterfaces preload, and crash-loop
stability are pre-existing package coverage. | `hermetic-default` |
Existing
`test/e2e-scenario/live/issue-2478-crash-loop-recovery.test.ts`,
`test/e2e-scenario/support-tests/e2e-recovery-helpers.test.ts`;
selective run `28186561267` job `issue-2478-crash-loop-recovery-vitest`
passed. No diff changes here. | existing / revalidated context |
| D2 | #5401 | Kimi remains a public-NVIDIA model/provider contract when
run in trusted selective CI, while retaining mock fallback for
local/untrusted validation. | `public-nvidia required` |
`.github/workflows/e2e-vitest-scenarios.yaml`,
`test/e2e-scenario/live/kimi-inference-compat.test.ts`,
`test/e2e-scenario/support-tests/kimi-inference-compat-helpers.test.ts`,
`test/e2e-script-workflow.test.ts` | covered / changed |
| D3 | #5413, #5625 | Kimi multiturn tool calls split `hostname; date;
uptime`, preserve tool-result flow, reject abandoned/continue traces,
and normalize final punctuation. | `public-nvidia required` with mock
fallback | `test/e2e-scenario/live/kimi-inference-compat-helpers.ts`
trajectory assertions; selective run `28190216767` job
`kimi-inference-compat-vitest` passed on the previous head; latest run
`28193896380` passed on `f36fef6da`. | covered / changed |
| D4 | #5406, #5412, #5760 | Scope-upgrade approval tolerates
preapproved / not-reproduced states, denies `operator.admin` leakage,
stays on gateway/no embedded fallback, and accepts whitespace-normalized
`42`; this is pre-existing package coverage. | `hosted-compatible
capable` | Existing
`test/e2e-scenario/live/issue-4462-scope-upgrade-approval.test.ts`;
selective run `28186561267` job
`issue-4462-scope-upgrade-approval-vitest` passed. No diff changes here.
| existing / revalidated context |

## Inference mode support
- Default mode for touched live target: Kimi `mock` unless workflow
selects `public-nvidia`.
- Real inference support preserved: yes for Kimi public NVIDIA; yes for
existing scope-upgrade hosted-compatible; not required for recovery.
- Modes validated in this PR: Kimi public NVIDIA via selective workflows
`28188683830`, `28190216767`; latest follow-up validation `28193896380`
is running for head `f36fef6da`. Kimi helper/mock behavior via local
support tests.
- Source-of-truth contract: `NEMOCLAW_E2E_INFERENCE_MODE` is the
canonical selector; absent selector defaults to mock for local/untrusted
validation; unknown explicit values now fail closed; legacy
`NEMOCLAW_KIMI_USE_MOCK=0` remains only as a temporary shell-lane
compatibility alias until shell retirement.
- Secret boundary: public Kimi workflow passes only `NVIDIA_API_KEY`;
helper probe envs are secret-free by default; raw public NVIDIA key
handoff is limited to onboard; sandbox `openclaw agent` now runs with a
secret-free env and uses the configured `nvidia-prod` route.

## Validation
- [x] `npx vitest run --project e2e-vitest-support
test/e2e-scenario/support-tests/kimi-inference-compat-helpers.test.ts`
- [x] `npx vitest run test/e2e-script-workflow.test.ts`
- [x] `npm run typecheck:cli`
- [x] `npm run test-conditionals:scan -- --top 25`
- [x] `npx prek run --all-files --stage pre-push --skip tsc-plugin
--skip tsc-js --skip tsc-cli --skip version-tag-sync --skip test-cli
--skip test-plugin --skip source-shape-test-budget --skip
test-file-size-budget --skip test-skills-yaml`
- [x] `git diff --check`
- [x] Kimi selective E2E / Vitest Scenarios on previous head:
https://github.com/NVIDIA/NemoClaw/actions/runs/28190216767
- [x] Kimi selective E2E / Vitest Scenarios after review-gap fixes:
https://github.com/NVIDIA/NemoClaw/actions/runs/28193896380
- [x] Existing recovery/scope rows revalidated in selective run:
https://github.com/NVIDIA/NemoClaw/actions/runs/28186561267
(`issue-2478-crash-loop-recovery-vitest` ✅,
`issue-4462-scope-upgrade-approval-vitest` ✅; Kimi in that stale run was
superseded)
- [ ] Local live mock Kimi: attempted but blocked by local Docker daemon
unavailable (`Cannot connect to the Docker daemon at
unix:///Users/jyaunches/.docker/run/docker.sock`). CI selective run is
the live validation path for this head.

## Follow-ups / waivers
- None.


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **New Features**
* Added support for running Kimi compatibility e2e checks in either mock
or public NVIDIA mode.
* The live scenario now adapts its setup, redaction, and traffic
validation based on the selected mode.

* **Bug Fixes**
* Improved handling of API key propagation so public NVIDIA runs use the
expected credentials without exposing secrets in other paths.

* **Tests**
* Added coverage for mode selection, API key validation, workflow
environment wiring, and the new public NVIDIA Vitest lane.

<!-- end of auto-generated comment: release notes by coderabbit.ai -->
@coderabbitai coderabbitai Bot mentioned this pull request Jun 26, 2026
5 of 6 tasks
Hadar301 pushed a commit to Hadar301/NemoClaw-OpenShift that referenced this pull request Jul 12, 2026
## Summary
Stabilizes the failing nightly hosted-inference E2E checks by making
live-answer assertions tolerant of harmless model-inserted whitespace
while keeping prompt-echo guards. Fixes the channels remove + rebuild
regression by preserving explicit empty messaging plans through rebuild
staging so token-backed channels are not rediscovered from the
environment.

## Related Issue
Fixes NVIDIA#5759

## Changes
- Preserve empty post-remove messaging plans in
`MessagingWorkflowPlanner.buildRebuildPlanFromSandboxEntry()` and stage
them into `NEMOCLAW_MESSAGING_PLAN_B64` during rebuild.
- Add regression coverage for final-channel removal rebuild plans and
rebuild staging of explicit empty messaging plans.
- Add shared E2E answer assertion helpers that normalize whitespace for
deterministic numeric/token replies.
- Update legacy bash E2Es and Vitest live replacements to use
whitespace-tolerant `42` and reply-token checks, including `full-e2e`,
`sandbox-operations`, `launchable-smoke`, `agent-turn-latency`,
`issue-4462`, GPU double-onboard, and OpenClaw TUI chat-correlation
paths.

## Type of Change
- [x] Code change (feature, bug fix, or refactor)
- [ ] Code change with doc updates
- [ ] Doc only (prose changes, no code sample modifications)
- [ ] Doc only (includes code sample changes)

## Verification
- [x] PR description includes the DCO sign-off declaration and every
commit appears as `Verified` in GitHub
- [x] Git hooks passed during commit and push, or `npx prek run
--from-ref main --to-ref HEAD` passes
- [x] Targeted tests pass for changed behavior
- [ ] Full `npm test` passes (broad runtime changes only)
- [x] Tests added or updated for new or changed behavior
- [x] No secrets, API keys, or credentials committed
- [ ] Docs updated for user-facing behavior changes
- [ ] `npm run docs` builds without warnings (doc changes only)
- [ ] Doc pages follow the [style
guide](https://github.com/NVIDIA/NemoClaw/blob/main/docs/CONTRIBUTING.md)
(doc changes only)
- [ ] New doc pages include SPDX header and frontmatter (new pages only)

Verification details:
- `npm run build:cli`
- `npx vitest run --project cli
test/helpers/e2e-answer-assertions.test.ts
test/openclaw-tui-chat-correlation.test.ts
src/lib/messaging/compiler/workflow-planner.test.ts
src/lib/actions/sandbox/rebuild-messaging-stage.test.ts`
- `npm run typecheck:cli`
- `npx prek run shellcheck --files test/e2e/lib/openclaw-json.sh
test/e2e/test-full-e2e.sh test/e2e/test-sandbox-operations.sh
test/e2e/test-launchable-smoke.sh
test/e2e/test-agent-turn-latency-e2e.sh
test/e2e/test-issue-4462-scope-upgrade-approval.sh`
- Commit and push hooks passed, including CLI tests and CLI TypeScript
pre-push checks.
- Documentation writer review: no docs changes needed because this is
E2E harness stabilization plus preserving the already-documented channel
remove/rebuild behavior.

---
<!-- DCO sign-off is required in this PR description, and every commit
must appear as Verified in GitHub. Run: git config user.name && git
config user.email -->
Signed-off-by: Carlos Villela <cvillela@nvidia.com>


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **Bug Fixes**
* Improved messaging rebuild planning so an explicitly empty channel set
is preserved (not cleared or re-derived).
* Made reply-token and integer `42` validations more robust to harmless
whitespace/newlines in streamed output.
* Enhanced chat-correlation checks to better tolerate formatting
variations.

* **Tests**
* Added/extended unit coverage for rebuild planning and “empty channels”
behavior.
* Introduced shared E2E assertion helpers and updated live/shell E2E
scenarios to use them.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Carlos Villela <cvillela@nvidia.com>
Hadar301 pushed a commit to Hadar301/NemoClaw-OpenShift that referenced this pull request Jul 12, 2026
## Summary
Ports the focused E2E stabilizers from `dep/openshell-v0.0.67` / PR
NVIDIA#5596 onto current `main` after PR NVIDIA#5760, without merging the full
OpenShell 0.0.67 branch.

This targets the full-main nightly failures from run 28172043426:
- `kimi-inference-compat-e2e` — relax live Kimi trajectory shape
expectations.
- `common-egress-agent-e2e` — tolerate wrapped reply tokens like
`REFER\nENCE_AGENT_OK`.
- `sessions-agents-cli-e2e` — keep sessions admin RPCs local/SDK-backed
and avoid multiline RPC args.

Also includes the small channel/remove rebuild staging stabilizer
carried by the shared matrix-stabilization commit.

## Validation
- Local push hooks could not fully run because this worktree is missing
local npm dependencies (`tsx`, `typescript`, Biome dependency `klaw`).
- Shellcheck/gitleaks/basic pre-commit checks passed before the
dependency-gated hooks failed.
- Focused nightly E2E dispatch is being run separately on this branch.

## Notes
- Does not port the full OpenShell 0.0.67 upgrade.
- Does not claim to fix `diagnostics-e2e` HTTP 403; that failure looked
infra/upstream/credential-like.

Signed-off-by: Julie Yaunches <jyaunches@nvidia.com>


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **Bug Fixes**
* Improved sandbox gateway RPC execution with pairing-aware retry, clear
retry/no-retry gating, and stricter handling of unsupported admin
methods.
* Added safer parsing and richer failure diagnostics with token
redaction in returned output and logged errors.
* **New Features**
* Enhanced gateway RPC results to include separate diagnostic output and
tightened admin method support via allowlisting.
* **Tests**
* Expanded Vitest coverage for gateway orchestration/output handling and
stream capture behavior.
* Strengthened OpenClaw text assertions, updated e2e token/PONG checks,
and relaxed Kimi validations for mock vs live.
  * Prevented Telegram env reuse after channel removal.
* **Chores**
* Added optional stdout/stderr stream capture controls for OpenShell
helpers.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Co-authored-by: Aaron Erickson <aerickson@nvidia.com>
Hadar301 pushed a commit to Hadar301/NemoClaw-OpenShift that referenced this pull request Jul 12, 2026
## Summary
Aligns the Vitest E2E lanes that receive `NVIDIA_INFERENCE_API_KEY` with
the hosted Inference Hub compatible-endpoint contract. This prevents
Build API `nvapi-` validation from rejecting the hosted-inference secret
before the affected scenarios reach their live assertions.

## Related Issue
Refs NVIDIA#5759

## Changes
- Configure `agent-turn-latency-vitest`, `launchable-smoke-vitest`, and
`openclaw-tui-chat-correlation-vitest` with hosted-compatible
provider/model env and `COMPATIBLE_API_KEY` sourced from
`NVIDIA_INFERENCE_API_KEY`.
- Teach the shared `cloud-openclaw` onboarding fixture to honor
`NEMOCLAW_E2E_USE_HOSTED_INFERENCE=1` by staging the compatible endpoint
provider and model.
- Update launchable smoke to skip `nvapi-` validation when running in
hosted-compatible mode and to expect the `compatible-endpoint` route.
- Add `--fresh`/`NEMOCLAW_FRESH=1` to agent-turn-latency install
attempts so retries do not trip over failed prior onboarding sessions.

## Type of Change
- [x] Code change (feature, bug fix, or refactor)
- [ ] Code change with doc updates
- [ ] Doc only (prose changes, no code sample modifications)
- [ ] Doc only (includes code sample changes)

## Verification
- [x] PR description includes the DCO sign-off declaration and every
commit appears as `Verified` in GitHub
- [x] Git hooks passed during commit and push, or `npx prek run
--from-ref main --to-ref HEAD` passes
- [x] Targeted tests pass for changed behavior
- [ ] Full `npm test` passes (broad runtime changes only)
- [x] Tests added or updated for new or changed behavior
- [x] No secrets, API keys, or credentials committed
- [ ] Docs updated for user-facing behavior changes
- [ ] `npm run docs` builds without warnings (doc changes only)
- [ ] Doc pages follow the [style
guide](https://github.com/NVIDIA/NemoClaw/blob/main/docs/CONTRIBUTING.md)
(doc changes only)
- [ ] New doc pages include SPDX header and frontmatter (new pages only)

Verification details:
- Synced with `origin/main` and confirmed the latest commit was NVIDIA#5760;
these Vitest lane hosted-key fixes were not already present.
- `npm run typecheck:cli`
- `npx vitest run --project cli
test/e2e-scenario/support-tests/e2e-scenarios-workflow.test.ts
test/e2e-script-workflow.test.ts
test/e2e-scenario/support-tests/hosted-inference.test.ts`
- `npx prek run check-yaml --files
.github/workflows/e2e-vitest-scenarios.yaml`
- Commit and push hooks passed, including CLI tests and CLI TypeScript
pre-push checks.
- Documentation writer review: no docs changes needed because this only
updates CI/live E2E harness behavior.

---
<!-- DCO sign-off is required in this PR description, and every commit
must appear as Verified in GitHub. Run: git config user.name && git
config user.email -->
Signed-off-by: Carlos Villela <cvillela@nvidia.com>


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **New Features**
* Added an opt-in “hosted inference compatibility” mode for end-to-end
onboarding and live scenarios, including hosted model/endpoint/provider
routing and onboarding environment setup.
* **Tests**
* Updated live smoke and helper flows to reflect hosted-compatible
behavior, including conditional API key checks and inference routing
assertions.
  * Improved test consistency with a “fresh” sandbox install option.
* **Chores**
* Updated CI end-to-end Vitest execution to run via the hosted NVIDIA
inference setup with secure API key injection and hosted model
configuration.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Carlos Villela <cvillela@nvidia.com>
Hadar301 pushed a commit to Hadar301/NemoClaw-OpenShift that referenced this pull request Jul 12, 2026
## Summary
Restore issue NVIDIA#5800 parity package `P0-B` for merged
onboard/rebuild/lifecycle bash-suite deltas only.

## Related Issues
Refs NVIDIA#5800
Refs NVIDIA#5098
Refs NVIDIA#5225
Refs NVIDIA#5487
Refs NVIDIA#5410
Refs NVIDIA#5760

## Scope gate
- Package: `P0-B — Onboard/rebuild/lifecycle parity`
- Included PRs all merged and touched `test/e2e`: yes — NVIDIA#5225, NVIDIA#5487,
NVIDIA#5410, NVIDIA#5760
- Out of scope: unmerged/non-bash PRs; shell lane retirement / PR NVIDIA#5756
cleanup

## Parity map
| ID | Source PR | Contract | Inference classification | Vitest
assertion / waiver | Status |
| --- | --- | --- | --- | --- | --- |
| B1 | NVIDIA#5225 | Persisted sandbox-entry gateway resolution is used by
lifecycle commands. | `none` | Existing
`src/lib/actions/sandbox/sandbox-gateway-routing.test.ts`,
`src/lib/onboard/gateway-binding.test.ts`,
`src/lib/onboard/sandbox-registration.test.ts` | covered |
| B2 | NVIDIA#5225 | Onboard repair/double-onboard failures include captured
onboard output diagnostics. | `none` | Existing live
`test/e2e-scenario/live/onboard-repair.test.ts`,
`test/e2e-scenario/live/double-onboard.test.ts`; diagnostics are
bash-runner-only verbosity and not a durable Vitest assertion. | waived
|
| B3 | NVIDIA#5487 | Plain `nemoclaw onboard` auto-detects an `in_progress`
session and resumes without `--resume`; `--fresh` suppresses
auto-resume. | `hosted-compatible capable` |
`test/e2e-scenario/live/onboard-resume.test.ts` Phase 3.5 mutates the
completed session to `in_progress`, asserts `(resume mode)` + cached
skips, then asserts `--fresh` fails at injected preflight without resume
banner. | covered |
| B4 | NVIDIA#5410 | Rebuild resumes messaging from `messaging.plan` rather
than legacy `providerCredentialHashes`; stale top-level provider hash
state is not used. | `none` |
`test/e2e-scenario/live/rebuild-hermes.test.ts` curated registry omits
`providerCredentialHashes`;
`src/lib/onboard/machine/handlers/sandbox.test.ts` refreshes
registry-plan credential hashes from env on rebuild resume. | covered |
| B5 | NVIDIA#5410 | Empty/staged rebuild messaging plan is preserved and
token-backed channels are not rediscovered. | `none` | Existing
`src/lib/actions/sandbox/rebuild-messaging-stage.test.ts` and
`src/lib/onboard/machine/handlers/sandbox.test.ts`. | covered |
| B6 | NVIDIA#5760 | Hosted-inference/messaging rebuild and live answer
assertions tolerate model whitespace around integer `42`. |
`hosted-compatible capable` | Existing
`test/helpers/e2e-answer-assertions.test.ts`, consumed by
`agent-turn-latency`, `full-e2e`, `launchable-smoke`, and
`sandbox-operations` live Vitests. | covered |
| B7 | NVIDIA#5760 | Stabilized hosted inference and messaging rebuild remain
validated by live full/sandbox/launchable/rebuild targets. |
`hosted-compatible capable` | Existing live targets remain unchanged;
local live execution blocked by Docker daemon unavailable. Selective
workflow required after PR opens. | follow-up |

## Inference mode support
- Default mode for touched live targets: `hosted-compatible capable` for
onboard-resume/rebuild-hermes/full/sandbox/launchable answer paths;
`none` for unit/process registry and gateway routing tests.
- Real inference support preserved: yes for existing hosted-compatible
live targets; no new inference adapter seam added here.
- Modes validated in this PR: local unit/process Vitests only; live
hosted-compatible validation needs GitHub runner/secrets because local
Docker daemon is unavailable.
- If not validated with real inference: local
`NEMOCLAW_RUN_E2E_SCENARIOS=1 npx vitest run --project
e2e-scenarios-live test/e2e-scenario/live/onboard-resume.test.ts
test/e2e-scenario/live/rebuild-hermes.test.ts` failed at prereq Docker
daemon check before scenario assertions.

## Validation
- [x] `git diff --check`
- [x] `npm test -- src/lib/onboard/machine/handlers/sandbox.test.ts
test/helpers/e2e-answer-assertions.test.ts
src/lib/actions/sandbox/sandbox-gateway-routing.test.ts
src/lib/onboard/entry-options.test.ts
src/lib/onboard/sandbox-registration.test.ts
src/lib/actions/sandbox/rebuild-messaging-stage.test.ts`
- [x] `npm run typecheck:cli`
- [x] `npm run build:cli`
- [x] `npm run test-size:check`
- [ ] hosted-compatible selective E2E workflow: pending PR / runner
dispatch

## Follow-ups / waivers
- Waiver B2: bash-only diagnostic verbosity from NVIDIA#5225 is not a durable
Vitest contract; existing live tests already preserve the functional
repair/double-onboard behavior.
- Follow-up B7: dispatch selective live Vitest scenarios on GitHub
runner with Docker + hosted inference secret after PR opens.


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **Tests**
* Added coverage for sandbox resume flows to verify updated Telegram
credentials are picked up when resuming a rebuild.
* Expanded end-to-end onboarding resume scenarios to confirm implicit
resume behavior, including skipped cached steps and fresh runs starting
from the expected point.
* Strengthened rebuild scenario checks to ensure curated registry
entries no longer include legacy credential-hash data.

<!-- end of auto-generated comment: release notes by coderabbit.ai -->
Hadar301 pushed a commit to Hadar301/NemoClaw-OpenShift that referenced this pull request Jul 12, 2026
## Summary
Restore the Kimi-specific issue NVIDIA#5800 parity work for package `P0-D`;
existing recovery and scope-upgrade package rows are explicitly mapped
as pre-existing coverage and revalidated context, but not changed
acceptance scope in this PR.

## Related Issues
Refs NVIDIA#5800
Refs NVIDIA#5098
Refs NVIDIA#5342
Refs NVIDIA#5401
Refs NVIDIA#5406
Refs NVIDIA#5412
Refs NVIDIA#5413
Refs NVIDIA#5625
Refs NVIDIA#5760

## Scope gate
- Package: `P0-D — Recovery, Kimi, and scope-upgrade parity`
- Included PRs all merged and touched `test/e2e`: yes
- Changed acceptance scope in this PR: Kimi public-NVIDIA/mock parity
(`D2`, `D3`)
- Existing package rows revalidated without diff changes: recovery
(`D1`) and scope-upgrade (`D4`)
- Out of scope: unmerged/non-bash PRs; shell lane retirement / PR NVIDIA#5756
cleanup

## Parity map
| ID | Source PR | Contract | Inference classification | Vitest
assertion / waiver | Status |
| --- | --- | --- | --- | --- | --- |
| D1 | NVIDIA#5342, NVIDIA#5401 | Recovery proxy env sourcing, missing proxy-env
warning, guard retention, ciao/networkInterfaces preload, and crash-loop
stability are pre-existing package coverage. | `hermetic-default` |
Existing
`test/e2e-scenario/live/issue-2478-crash-loop-recovery.test.ts`,
`test/e2e-scenario/support-tests/e2e-recovery-helpers.test.ts`;
selective run `28186561267` job `issue-2478-crash-loop-recovery-vitest`
passed. No diff changes here. | existing / revalidated context |
| D2 | NVIDIA#5401 | Kimi remains a public-NVIDIA model/provider contract when
run in trusted selective CI, while retaining mock fallback for
local/untrusted validation. | `public-nvidia required` |
`.github/workflows/e2e-vitest-scenarios.yaml`,
`test/e2e-scenario/live/kimi-inference-compat.test.ts`,
`test/e2e-scenario/support-tests/kimi-inference-compat-helpers.test.ts`,
`test/e2e-script-workflow.test.ts` | covered / changed |
| D3 | NVIDIA#5413, NVIDIA#5625 | Kimi multiturn tool calls split `hostname; date;
uptime`, preserve tool-result flow, reject abandoned/continue traces,
and normalize final punctuation. | `public-nvidia required` with mock
fallback | `test/e2e-scenario/live/kimi-inference-compat-helpers.ts`
trajectory assertions; selective run `28190216767` job
`kimi-inference-compat-vitest` passed on the previous head; latest run
`28193896380` passed on `f36fef6da`. | covered / changed |
| D4 | NVIDIA#5406, NVIDIA#5412, NVIDIA#5760 | Scope-upgrade approval tolerates
preapproved / not-reproduced states, denies `operator.admin` leakage,
stays on gateway/no embedded fallback, and accepts whitespace-normalized
`42`; this is pre-existing package coverage. | `hosted-compatible
capable` | Existing
`test/e2e-scenario/live/issue-4462-scope-upgrade-approval.test.ts`;
selective run `28186561267` job
`issue-4462-scope-upgrade-approval-vitest` passed. No diff changes here.
| existing / revalidated context |

## Inference mode support
- Default mode for touched live target: Kimi `mock` unless workflow
selects `public-nvidia`.
- Real inference support preserved: yes for Kimi public NVIDIA; yes for
existing scope-upgrade hosted-compatible; not required for recovery.
- Modes validated in this PR: Kimi public NVIDIA via selective workflows
`28188683830`, `28190216767`; latest follow-up validation `28193896380`
is running for head `f36fef6da`. Kimi helper/mock behavior via local
support tests.
- Source-of-truth contract: `NEMOCLAW_E2E_INFERENCE_MODE` is the
canonical selector; absent selector defaults to mock for local/untrusted
validation; unknown explicit values now fail closed; legacy
`NEMOCLAW_KIMI_USE_MOCK=0` remains only as a temporary shell-lane
compatibility alias until shell retirement.
- Secret boundary: public Kimi workflow passes only `NVIDIA_API_KEY`;
helper probe envs are secret-free by default; raw public NVIDIA key
handoff is limited to onboard; sandbox `openclaw agent` now runs with a
secret-free env and uses the configured `nvidia-prod` route.

## Validation
- [x] `npx vitest run --project e2e-vitest-support
test/e2e-scenario/support-tests/kimi-inference-compat-helpers.test.ts`
- [x] `npx vitest run test/e2e-script-workflow.test.ts`
- [x] `npm run typecheck:cli`
- [x] `npm run test-conditionals:scan -- --top 25`
- [x] `npx prek run --all-files --stage pre-push --skip tsc-plugin
--skip tsc-js --skip tsc-cli --skip version-tag-sync --skip test-cli
--skip test-plugin --skip source-shape-test-budget --skip
test-file-size-budget --skip test-skills-yaml`
- [x] `git diff --check`
- [x] Kimi selective E2E / Vitest Scenarios on previous head:
https://github.com/NVIDIA/NemoClaw/actions/runs/28190216767
- [x] Kimi selective E2E / Vitest Scenarios after review-gap fixes:
https://github.com/NVIDIA/NemoClaw/actions/runs/28193896380
- [x] Existing recovery/scope rows revalidated in selective run:
https://github.com/NVIDIA/NemoClaw/actions/runs/28186561267
(`issue-2478-crash-loop-recovery-vitest` ✅,
`issue-4462-scope-upgrade-approval-vitest` ✅; Kimi in that stale run was
superseded)
- [ ] Local live mock Kimi: attempted but blocked by local Docker daemon
unavailable (`Cannot connect to the Docker daemon at
unix:///Users/jyaunches/.docker/run/docker.sock`). CI selective run is
the live validation path for this head.

## Follow-ups / waivers
- None.


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **New Features**
* Added support for running Kimi compatibility e2e checks in either mock
or public NVIDIA mode.
* The live scenario now adapts its setup, redaction, and traffic
validation based on the selected mode.

* **Bug Fixes**
* Improved handling of API key propagation so public NVIDIA runs use the
expected credentials without exposing secrets in other paths.

* **Tests**
* Added coverage for mode selection, API key validation, workflow
environment wiring, and the new public NVIDIA Vitest lane.

<!-- end of auto-generated comment: release notes by coderabbit.ai -->
@wscurran wscurran added area: e2e End-to-end tests, nightly failures, or validation infrastructure area: messaging Messaging channels, bridges, manifests, or channel lifecycle bug-fix PR fixes a bug or regression labels Aug 1, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

area: e2e End-to-end tests, nightly failures, or validation infrastructure area: messaging Messaging channels, bridges, manifests, or channel lifecycle bug-fix PR fixes a bug or regression

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Nightly E2E failures: brittle hosted-model assertions and channels remove rebuild regression

4 participants