fix(e2e): stabilize hosted inference and messaging rebuild - #5760
Conversation
Signed-off-by: Carlos Villela <cvillela@nvidia.com>
|
No actionable comments were generated in the recent review. 🎉 ℹ️ Recent review info⚙️ Run configurationConfiguration used: Path: .coderabbit.yaml Review profile: CHILL Plan: Enterprise Run ID: 📒 Files selected for processing (1)
🚧 Files skipped from review as they are similar to previous changes (1)
📝 WalkthroughWalkthroughThe PR preserves explicit empty messaging rebuild plans after channel removal and replaces exact 42/reply-token matching with whitespace-tolerant helpers across Vitest and shell E2E scenarios. ChangesMessaging rebuild and E2E assertions
Possibly related PRs
Suggested labels
Suggested reviewers
Estimated code review effort🎯 3 (Moderate) | ⏱️ ~30 minutes Poem
🚥 Pre-merge checks | ✅ 4 | ❌ 1❌ Failed checks (1 warning)
✅ Passed checks (4 passed)
✏️ Tip: You can configure your own custom pre-merge checks in the settings. ✨ Finishing Touches📝 Generate docstrings
🧪 Generate unit tests (beta)
Comment |
Code Coverage OverviewLanguages: TypeScript TypeScript / code-coverage/pluginThe overall coverage in the Show a code coverage summary of the most covered files.
TypeScript / code-coverage/cliThe overall coverage in the Show a code coverage summary of the most covered files.
Updated |
PR Review Advisor — No blocking findingsMerge posture: No blocking advisor findings Action checklist
Test follow-ups to resolve or justifyIf these cover changed behavior, prefer adding them in this PR; otherwise state why existing coverage is enough or link the follow-up.
This is an automated, non-binding review; it still expects maintainers and agents to respond to each required or warning item. Treat suggestions as current-PR improvements when they touch changed code; defer only with maintainer rationale or a linked follow-up. A human maintainer must make the final merge decision. |
E2E Advisor RecommendationRequired E2E: Dispatch hint: Full advisor summaryE2E Recommendation AdvisorBase: Required E2E
Optional E2E
New E2E recommendations
Dispatch hint
|
Vitest E2E Scenario RecommendationRequired Vitest E2E scenarios: Dispatch required Vitest E2E scenarios:
Full Vitest E2E advisor summaryVitest E2E Scenario AdvisorBase: Required Vitest E2E scenarios
Optional Vitest E2E scenarios
Relevant changed files
|
There was a problem hiding this comment.
Actionable comments posted: 1
🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.
Inline comments:
In `@test/e2e-scenario/live/issue-4462-scope-upgrade-approval.test.ts`:
- Around line 122-124: The contains_integer_42 helper is using a brittle tr |
grep pipeline that can fail under set -euo pipefail. Update contains_integer_42
to follow the captured-string approach used in test/e2e/lib/openclaw-json.sh, so
it first reads the input into a variable, strips whitespace from that value, and
then performs the integer 42 match without relying on a pipeline.
🪄 Autofix (Beta)
Fix all unresolved CodeRabbit comments on this PR:
- Push a commit to this branch (recommended)
- Create a new PR with the fixes
ℹ️ Review info
⚙️ Run configuration
Configuration used: Path: .coderabbit.yaml
Review profile: CHILL
Plan: Enterprise
Run ID: fb6ff966-d8a8-4ea6-a01c-845bbebb1a77
📒 Files selected for processing (19)
src/lib/actions/sandbox/rebuild-messaging-stage.test.tssrc/lib/messaging/compiler/workflow-planner.test.tssrc/lib/messaging/compiler/workflow-planner.tstest/e2e-scenario/live/agent-turn-latency.test.tstest/e2e-scenario/live/full-e2e.test.tstest/e2e-scenario/live/gpu-double-onboard.test.tstest/e2e-scenario/live/issue-4462-scope-upgrade-approval.test.tstest/e2e-scenario/live/launchable-smoke.test.tstest/e2e-scenario/live/openclaw-tui-chat-correlation.test.tstest/e2e-scenario/live/sandbox-operations.test.tstest/e2e/lib/openclaw-json.shtest/e2e/test-agent-turn-latency-e2e.shtest/e2e/test-full-e2e.shtest/e2e/test-issue-4462-scope-upgrade-approval.shtest/e2e/test-launchable-smoke.shtest/e2e/test-sandbox-operations.shtest/helpers/e2e-answer-assertions.test.tstest/helpers/e2e-answer-assertions.tstest/openclaw-tui-chat-correlation.test.ts
Signed-off-by: Carlos Villela <cvillela@nvidia.com>
Vitest E2E Scenario Results — ✅ All requested jobs passedRun: 28144691365
|
Vitest E2E Scenario Results — ❌ Some jobs failedRun: 28144742533
|
|
E2E follow-up for the advisor items:
So the targeted runtime proof for the real messaging regression is green, and the additional live assertion lanes that failed did so before exercising the whitespace-normalized assertion changes. |
## Summary Ports the focused E2E stabilizers from `dep/openshell-v0.0.67` / PR #5596 onto current `main` after PR #5760, without merging the full OpenShell 0.0.67 branch. This targets the full-main nightly failures from run 28172043426: - `kimi-inference-compat-e2e` — relax live Kimi trajectory shape expectations. - `common-egress-agent-e2e` — tolerate wrapped reply tokens like `REFER\nENCE_AGENT_OK`. - `sessions-agents-cli-e2e` — keep sessions admin RPCs local/SDK-backed and avoid multiline RPC args. Also includes the small channel/remove rebuild staging stabilizer carried by the shared matrix-stabilization commit. ## Validation - Local push hooks could not fully run because this worktree is missing local npm dependencies (`tsx`, `typescript`, Biome dependency `klaw`). - Shellcheck/gitleaks/basic pre-commit checks passed before the dependency-gated hooks failed. - Focused nightly E2E dispatch is being run separately on this branch. ## Notes - Does not port the full OpenShell 0.0.67 upgrade. - Does not claim to fix `diagnostics-e2e` HTTP 403; that failure looked infra/upstream/credential-like. Signed-off-by: Julie Yaunches <jyaunches@nvidia.com> <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Bug Fixes** * Improved sandbox gateway RPC execution with pairing-aware retry, clear retry/no-retry gating, and stricter handling of unsupported admin methods. * Added safer parsing and richer failure diagnostics with token redaction in returned output and logged errors. * **New Features** * Enhanced gateway RPC results to include separate diagnostic output and tightened admin method support via allowlisting. * **Tests** * Expanded Vitest coverage for gateway orchestration/output handling and stream capture behavior. * Strengthened OpenClaw text assertions, updated e2e token/PONG checks, and relaxed Kimi validations for mock vs live. * Prevented Telegram env reuse after channel removal. * **Chores** * Added optional stdout/stderr stream capture controls for OpenShell helpers. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Co-authored-by: Aaron Erickson <aerickson@nvidia.com>
## Summary Aligns the Vitest E2E lanes that receive `NVIDIA_INFERENCE_API_KEY` with the hosted Inference Hub compatible-endpoint contract. This prevents Build API `nvapi-` validation from rejecting the hosted-inference secret before the affected scenarios reach their live assertions. ## Related Issue Refs #5759 ## Changes - Configure `agent-turn-latency-vitest`, `launchable-smoke-vitest`, and `openclaw-tui-chat-correlation-vitest` with hosted-compatible provider/model env and `COMPATIBLE_API_KEY` sourced from `NVIDIA_INFERENCE_API_KEY`. - Teach the shared `cloud-openclaw` onboarding fixture to honor `NEMOCLAW_E2E_USE_HOSTED_INFERENCE=1` by staging the compatible endpoint provider and model. - Update launchable smoke to skip `nvapi-` validation when running in hosted-compatible mode and to expect the `compatible-endpoint` route. - Add `--fresh`/`NEMOCLAW_FRESH=1` to agent-turn-latency install attempts so retries do not trip over failed prior onboarding sessions. ## Type of Change - [x] Code change (feature, bug fix, or refactor) - [ ] Code change with doc updates - [ ] Doc only (prose changes, no code sample modifications) - [ ] Doc only (includes code sample changes) ## Verification - [x] PR description includes the DCO sign-off declaration and every commit appears as `Verified` in GitHub - [x] Git hooks passed during commit and push, or `npx prek run --from-ref main --to-ref HEAD` passes - [x] Targeted tests pass for changed behavior - [ ] Full `npm test` passes (broad runtime changes only) - [x] Tests added or updated for new or changed behavior - [x] No secrets, API keys, or credentials committed - [ ] Docs updated for user-facing behavior changes - [ ] `npm run docs` builds without warnings (doc changes only) - [ ] Doc pages follow the [style guide](https://github.com/NVIDIA/NemoClaw/blob/main/docs/CONTRIBUTING.md) (doc changes only) - [ ] New doc pages include SPDX header and frontmatter (new pages only) Verification details: - Synced with `origin/main` and confirmed the latest commit was #5760; these Vitest lane hosted-key fixes were not already present. - `npm run typecheck:cli` - `npx vitest run --project cli test/e2e-scenario/support-tests/e2e-scenarios-workflow.test.ts test/e2e-script-workflow.test.ts test/e2e-scenario/support-tests/hosted-inference.test.ts` - `npx prek run check-yaml --files .github/workflows/e2e-vitest-scenarios.yaml` - Commit and push hooks passed, including CLI tests and CLI TypeScript pre-push checks. - Documentation writer review: no docs changes needed because this only updates CI/live E2E harness behavior. --- <!-- DCO sign-off is required in this PR description, and every commit must appear as Verified in GitHub. Run: git config user.name && git config user.email --> Signed-off-by: Carlos Villela <cvillela@nvidia.com> <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **New Features** * Added an opt-in “hosted inference compatibility” mode for end-to-end onboarding and live scenarios, including hosted model/endpoint/provider routing and onboarding environment setup. * **Tests** * Updated live smoke and helper flows to reflect hosted-compatible behavior, including conditional API key checks and inference routing assertions. * Improved test consistency with a “fresh” sandbox install option. * **Chores** * Updated CI end-to-end Vitest execution to run via the hosted NVIDIA inference setup with secure API key injection and hosted model configuration. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Carlos Villela <cvillela@nvidia.com>
## Summary Restore issue #5800 parity package `P0-B` for merged onboard/rebuild/lifecycle bash-suite deltas only. ## Related Issues Refs #5800 Refs #5098 Refs #5225 Refs #5487 Refs #5410 Refs #5760 ## Scope gate - Package: `P0-B — Onboard/rebuild/lifecycle parity` - Included PRs all merged and touched `test/e2e`: yes — #5225, #5487, #5410, #5760 - Out of scope: unmerged/non-bash PRs; shell lane retirement / PR #5756 cleanup ## Parity map | ID | Source PR | Contract | Inference classification | Vitest assertion / waiver | Status | | --- | --- | --- | --- | --- | --- | | B1 | #5225 | Persisted sandbox-entry gateway resolution is used by lifecycle commands. | `none` | Existing `src/lib/actions/sandbox/sandbox-gateway-routing.test.ts`, `src/lib/onboard/gateway-binding.test.ts`, `src/lib/onboard/sandbox-registration.test.ts` | covered | | B2 | #5225 | Onboard repair/double-onboard failures include captured onboard output diagnostics. | `none` | Existing live `test/e2e-scenario/live/onboard-repair.test.ts`, `test/e2e-scenario/live/double-onboard.test.ts`; diagnostics are bash-runner-only verbosity and not a durable Vitest assertion. | waived | | B3 | #5487 | Plain `nemoclaw onboard` auto-detects an `in_progress` session and resumes without `--resume`; `--fresh` suppresses auto-resume. | `hosted-compatible capable` | `test/e2e-scenario/live/onboard-resume.test.ts` Phase 3.5 mutates the completed session to `in_progress`, asserts `(resume mode)` + cached skips, then asserts `--fresh` fails at injected preflight without resume banner. | covered | | B4 | #5410 | Rebuild resumes messaging from `messaging.plan` rather than legacy `providerCredentialHashes`; stale top-level provider hash state is not used. | `none` | `test/e2e-scenario/live/rebuild-hermes.test.ts` curated registry omits `providerCredentialHashes`; `src/lib/onboard/machine/handlers/sandbox.test.ts` refreshes registry-plan credential hashes from env on rebuild resume. | covered | | B5 | #5410 | Empty/staged rebuild messaging plan is preserved and token-backed channels are not rediscovered. | `none` | Existing `src/lib/actions/sandbox/rebuild-messaging-stage.test.ts` and `src/lib/onboard/machine/handlers/sandbox.test.ts`. | covered | | B6 | #5760 | Hosted-inference/messaging rebuild and live answer assertions tolerate model whitespace around integer `42`. | `hosted-compatible capable` | Existing `test/helpers/e2e-answer-assertions.test.ts`, consumed by `agent-turn-latency`, `full-e2e`, `launchable-smoke`, and `sandbox-operations` live Vitests. | covered | | B7 | #5760 | Stabilized hosted inference and messaging rebuild remain validated by live full/sandbox/launchable/rebuild targets. | `hosted-compatible capable` | Existing live targets remain unchanged; local live execution blocked by Docker daemon unavailable. Selective workflow required after PR opens. | follow-up | ## Inference mode support - Default mode for touched live targets: `hosted-compatible capable` for onboard-resume/rebuild-hermes/full/sandbox/launchable answer paths; `none` for unit/process registry and gateway routing tests. - Real inference support preserved: yes for existing hosted-compatible live targets; no new inference adapter seam added here. - Modes validated in this PR: local unit/process Vitests only; live hosted-compatible validation needs GitHub runner/secrets because local Docker daemon is unavailable. - If not validated with real inference: local `NEMOCLAW_RUN_E2E_SCENARIOS=1 npx vitest run --project e2e-scenarios-live test/e2e-scenario/live/onboard-resume.test.ts test/e2e-scenario/live/rebuild-hermes.test.ts` failed at prereq Docker daemon check before scenario assertions. ## Validation - [x] `git diff --check` - [x] `npm test -- src/lib/onboard/machine/handlers/sandbox.test.ts test/helpers/e2e-answer-assertions.test.ts src/lib/actions/sandbox/sandbox-gateway-routing.test.ts src/lib/onboard/entry-options.test.ts src/lib/onboard/sandbox-registration.test.ts src/lib/actions/sandbox/rebuild-messaging-stage.test.ts` - [x] `npm run typecheck:cli` - [x] `npm run build:cli` - [x] `npm run test-size:check` - [ ] hosted-compatible selective E2E workflow: pending PR / runner dispatch ## Follow-ups / waivers - Waiver B2: bash-only diagnostic verbosity from #5225 is not a durable Vitest contract; existing live tests already preserve the functional repair/double-onboard behavior. - Follow-up B7: dispatch selective live Vitest scenarios on GitHub runner with Docker + hosted inference secret after PR opens. <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Tests** * Added coverage for sandbox resume flows to verify updated Telegram credentials are picked up when resuming a rebuild. * Expanded end-to-end onboarding resume scenarios to confirm implicit resume behavior, including skipped cached steps and fresh runs starting from the expected point. * Strengthened rebuild scenario checks to ensure curated registry entries no longer include legacy credential-hash data. <!-- end of auto-generated comment: release notes by coderabbit.ai -->
## Summary Restore the Kimi-specific issue #5800 parity work for package `P0-D`; existing recovery and scope-upgrade package rows are explicitly mapped as pre-existing coverage and revalidated context, but not changed acceptance scope in this PR. ## Related Issues Refs #5800 Refs #5098 Refs #5342 Refs #5401 Refs #5406 Refs #5412 Refs #5413 Refs #5625 Refs #5760 ## Scope gate - Package: `P0-D — Recovery, Kimi, and scope-upgrade parity` - Included PRs all merged and touched `test/e2e`: yes - Changed acceptance scope in this PR: Kimi public-NVIDIA/mock parity (`D2`, `D3`) - Existing package rows revalidated without diff changes: recovery (`D1`) and scope-upgrade (`D4`) - Out of scope: unmerged/non-bash PRs; shell lane retirement / PR #5756 cleanup ## Parity map | ID | Source PR | Contract | Inference classification | Vitest assertion / waiver | Status | | --- | --- | --- | --- | --- | --- | | D1 | #5342, #5401 | Recovery proxy env sourcing, missing proxy-env warning, guard retention, ciao/networkInterfaces preload, and crash-loop stability are pre-existing package coverage. | `hermetic-default` | Existing `test/e2e-scenario/live/issue-2478-crash-loop-recovery.test.ts`, `test/e2e-scenario/support-tests/e2e-recovery-helpers.test.ts`; selective run `28186561267` job `issue-2478-crash-loop-recovery-vitest` passed. No diff changes here. | existing / revalidated context | | D2 | #5401 | Kimi remains a public-NVIDIA model/provider contract when run in trusted selective CI, while retaining mock fallback for local/untrusted validation. | `public-nvidia required` | `.github/workflows/e2e-vitest-scenarios.yaml`, `test/e2e-scenario/live/kimi-inference-compat.test.ts`, `test/e2e-scenario/support-tests/kimi-inference-compat-helpers.test.ts`, `test/e2e-script-workflow.test.ts` | covered / changed | | D3 | #5413, #5625 | Kimi multiturn tool calls split `hostname; date; uptime`, preserve tool-result flow, reject abandoned/continue traces, and normalize final punctuation. | `public-nvidia required` with mock fallback | `test/e2e-scenario/live/kimi-inference-compat-helpers.ts` trajectory assertions; selective run `28190216767` job `kimi-inference-compat-vitest` passed on the previous head; latest run `28193896380` passed on `f36fef6da`. | covered / changed | | D4 | #5406, #5412, #5760 | Scope-upgrade approval tolerates preapproved / not-reproduced states, denies `operator.admin` leakage, stays on gateway/no embedded fallback, and accepts whitespace-normalized `42`; this is pre-existing package coverage. | `hosted-compatible capable` | Existing `test/e2e-scenario/live/issue-4462-scope-upgrade-approval.test.ts`; selective run `28186561267` job `issue-4462-scope-upgrade-approval-vitest` passed. No diff changes here. | existing / revalidated context | ## Inference mode support - Default mode for touched live target: Kimi `mock` unless workflow selects `public-nvidia`. - Real inference support preserved: yes for Kimi public NVIDIA; yes for existing scope-upgrade hosted-compatible; not required for recovery. - Modes validated in this PR: Kimi public NVIDIA via selective workflows `28188683830`, `28190216767`; latest follow-up validation `28193896380` is running for head `f36fef6da`. Kimi helper/mock behavior via local support tests. - Source-of-truth contract: `NEMOCLAW_E2E_INFERENCE_MODE` is the canonical selector; absent selector defaults to mock for local/untrusted validation; unknown explicit values now fail closed; legacy `NEMOCLAW_KIMI_USE_MOCK=0` remains only as a temporary shell-lane compatibility alias until shell retirement. - Secret boundary: public Kimi workflow passes only `NVIDIA_API_KEY`; helper probe envs are secret-free by default; raw public NVIDIA key handoff is limited to onboard; sandbox `openclaw agent` now runs with a secret-free env and uses the configured `nvidia-prod` route. ## Validation - [x] `npx vitest run --project e2e-vitest-support test/e2e-scenario/support-tests/kimi-inference-compat-helpers.test.ts` - [x] `npx vitest run test/e2e-script-workflow.test.ts` - [x] `npm run typecheck:cli` - [x] `npm run test-conditionals:scan -- --top 25` - [x] `npx prek run --all-files --stage pre-push --skip tsc-plugin --skip tsc-js --skip tsc-cli --skip version-tag-sync --skip test-cli --skip test-plugin --skip source-shape-test-budget --skip test-file-size-budget --skip test-skills-yaml` - [x] `git diff --check` - [x] Kimi selective E2E / Vitest Scenarios on previous head: https://github.com/NVIDIA/NemoClaw/actions/runs/28190216767 - [x] Kimi selective E2E / Vitest Scenarios after review-gap fixes: https://github.com/NVIDIA/NemoClaw/actions/runs/28193896380 - [x] Existing recovery/scope rows revalidated in selective run: https://github.com/NVIDIA/NemoClaw/actions/runs/28186561267 (`issue-2478-crash-loop-recovery-vitest` ✅, `issue-4462-scope-upgrade-approval-vitest` ✅; Kimi in that stale run was superseded) - [ ] Local live mock Kimi: attempted but blocked by local Docker daemon unavailable (`Cannot connect to the Docker daemon at unix:///Users/jyaunches/.docker/run/docker.sock`). CI selective run is the live validation path for this head. ## Follow-ups / waivers - None. <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **New Features** * Added support for running Kimi compatibility e2e checks in either mock or public NVIDIA mode. * The live scenario now adapts its setup, redaction, and traffic validation based on the selected mode. * **Bug Fixes** * Improved handling of API key propagation so public NVIDIA runs use the expected credentials without exposing secrets in other paths. * **Tests** * Added coverage for mode selection, API key validation, workflow environment wiring, and the new public NVIDIA Vitest lane. <!-- end of auto-generated comment: release notes by coderabbit.ai -->
## Summary Stabilizes the failing nightly hosted-inference E2E checks by making live-answer assertions tolerant of harmless model-inserted whitespace while keeping prompt-echo guards. Fixes the channels remove + rebuild regression by preserving explicit empty messaging plans through rebuild staging so token-backed channels are not rediscovered from the environment. ## Related Issue Fixes NVIDIA#5759 ## Changes - Preserve empty post-remove messaging plans in `MessagingWorkflowPlanner.buildRebuildPlanFromSandboxEntry()` and stage them into `NEMOCLAW_MESSAGING_PLAN_B64` during rebuild. - Add regression coverage for final-channel removal rebuild plans and rebuild staging of explicit empty messaging plans. - Add shared E2E answer assertion helpers that normalize whitespace for deterministic numeric/token replies. - Update legacy bash E2Es and Vitest live replacements to use whitespace-tolerant `42` and reply-token checks, including `full-e2e`, `sandbox-operations`, `launchable-smoke`, `agent-turn-latency`, `issue-4462`, GPU double-onboard, and OpenClaw TUI chat-correlation paths. ## Type of Change - [x] Code change (feature, bug fix, or refactor) - [ ] Code change with doc updates - [ ] Doc only (prose changes, no code sample modifications) - [ ] Doc only (includes code sample changes) ## Verification - [x] PR description includes the DCO sign-off declaration and every commit appears as `Verified` in GitHub - [x] Git hooks passed during commit and push, or `npx prek run --from-ref main --to-ref HEAD` passes - [x] Targeted tests pass for changed behavior - [ ] Full `npm test` passes (broad runtime changes only) - [x] Tests added or updated for new or changed behavior - [x] No secrets, API keys, or credentials committed - [ ] Docs updated for user-facing behavior changes - [ ] `npm run docs` builds without warnings (doc changes only) - [ ] Doc pages follow the [style guide](https://github.com/NVIDIA/NemoClaw/blob/main/docs/CONTRIBUTING.md) (doc changes only) - [ ] New doc pages include SPDX header and frontmatter (new pages only) Verification details: - `npm run build:cli` - `npx vitest run --project cli test/helpers/e2e-answer-assertions.test.ts test/openclaw-tui-chat-correlation.test.ts src/lib/messaging/compiler/workflow-planner.test.ts src/lib/actions/sandbox/rebuild-messaging-stage.test.ts` - `npm run typecheck:cli` - `npx prek run shellcheck --files test/e2e/lib/openclaw-json.sh test/e2e/test-full-e2e.sh test/e2e/test-sandbox-operations.sh test/e2e/test-launchable-smoke.sh test/e2e/test-agent-turn-latency-e2e.sh test/e2e/test-issue-4462-scope-upgrade-approval.sh` - Commit and push hooks passed, including CLI tests and CLI TypeScript pre-push checks. - Documentation writer review: no docs changes needed because this is E2E harness stabilization plus preserving the already-documented channel remove/rebuild behavior. --- <!-- DCO sign-off is required in this PR description, and every commit must appear as Verified in GitHub. Run: git config user.name && git config user.email --> Signed-off-by: Carlos Villela <cvillela@nvidia.com> <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Bug Fixes** * Improved messaging rebuild planning so an explicitly empty channel set is preserved (not cleared or re-derived). * Made reply-token and integer `42` validations more robust to harmless whitespace/newlines in streamed output. * Enhanced chat-correlation checks to better tolerate formatting variations. * **Tests** * Added/extended unit coverage for rebuild planning and “empty channels” behavior. * Introduced shared E2E assertion helpers and updated live/shell E2E scenarios to use them. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Carlos Villela <cvillela@nvidia.com>
## Summary Ports the focused E2E stabilizers from `dep/openshell-v0.0.67` / PR NVIDIA#5596 onto current `main` after PR NVIDIA#5760, without merging the full OpenShell 0.0.67 branch. This targets the full-main nightly failures from run 28172043426: - `kimi-inference-compat-e2e` — relax live Kimi trajectory shape expectations. - `common-egress-agent-e2e` — tolerate wrapped reply tokens like `REFER\nENCE_AGENT_OK`. - `sessions-agents-cli-e2e` — keep sessions admin RPCs local/SDK-backed and avoid multiline RPC args. Also includes the small channel/remove rebuild staging stabilizer carried by the shared matrix-stabilization commit. ## Validation - Local push hooks could not fully run because this worktree is missing local npm dependencies (`tsx`, `typescript`, Biome dependency `klaw`). - Shellcheck/gitleaks/basic pre-commit checks passed before the dependency-gated hooks failed. - Focused nightly E2E dispatch is being run separately on this branch. ## Notes - Does not port the full OpenShell 0.0.67 upgrade. - Does not claim to fix `diagnostics-e2e` HTTP 403; that failure looked infra/upstream/credential-like. Signed-off-by: Julie Yaunches <jyaunches@nvidia.com> <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Bug Fixes** * Improved sandbox gateway RPC execution with pairing-aware retry, clear retry/no-retry gating, and stricter handling of unsupported admin methods. * Added safer parsing and richer failure diagnostics with token redaction in returned output and logged errors. * **New Features** * Enhanced gateway RPC results to include separate diagnostic output and tightened admin method support via allowlisting. * **Tests** * Expanded Vitest coverage for gateway orchestration/output handling and stream capture behavior. * Strengthened OpenClaw text assertions, updated e2e token/PONG checks, and relaxed Kimi validations for mock vs live. * Prevented Telegram env reuse after channel removal. * **Chores** * Added optional stdout/stderr stream capture controls for OpenShell helpers. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Co-authored-by: Aaron Erickson <aerickson@nvidia.com>
## Summary Aligns the Vitest E2E lanes that receive `NVIDIA_INFERENCE_API_KEY` with the hosted Inference Hub compatible-endpoint contract. This prevents Build API `nvapi-` validation from rejecting the hosted-inference secret before the affected scenarios reach their live assertions. ## Related Issue Refs NVIDIA#5759 ## Changes - Configure `agent-turn-latency-vitest`, `launchable-smoke-vitest`, and `openclaw-tui-chat-correlation-vitest` with hosted-compatible provider/model env and `COMPATIBLE_API_KEY` sourced from `NVIDIA_INFERENCE_API_KEY`. - Teach the shared `cloud-openclaw` onboarding fixture to honor `NEMOCLAW_E2E_USE_HOSTED_INFERENCE=1` by staging the compatible endpoint provider and model. - Update launchable smoke to skip `nvapi-` validation when running in hosted-compatible mode and to expect the `compatible-endpoint` route. - Add `--fresh`/`NEMOCLAW_FRESH=1` to agent-turn-latency install attempts so retries do not trip over failed prior onboarding sessions. ## Type of Change - [x] Code change (feature, bug fix, or refactor) - [ ] Code change with doc updates - [ ] Doc only (prose changes, no code sample modifications) - [ ] Doc only (includes code sample changes) ## Verification - [x] PR description includes the DCO sign-off declaration and every commit appears as `Verified` in GitHub - [x] Git hooks passed during commit and push, or `npx prek run --from-ref main --to-ref HEAD` passes - [x] Targeted tests pass for changed behavior - [ ] Full `npm test` passes (broad runtime changes only) - [x] Tests added or updated for new or changed behavior - [x] No secrets, API keys, or credentials committed - [ ] Docs updated for user-facing behavior changes - [ ] `npm run docs` builds without warnings (doc changes only) - [ ] Doc pages follow the [style guide](https://github.com/NVIDIA/NemoClaw/blob/main/docs/CONTRIBUTING.md) (doc changes only) - [ ] New doc pages include SPDX header and frontmatter (new pages only) Verification details: - Synced with `origin/main` and confirmed the latest commit was NVIDIA#5760; these Vitest lane hosted-key fixes were not already present. - `npm run typecheck:cli` - `npx vitest run --project cli test/e2e-scenario/support-tests/e2e-scenarios-workflow.test.ts test/e2e-script-workflow.test.ts test/e2e-scenario/support-tests/hosted-inference.test.ts` - `npx prek run check-yaml --files .github/workflows/e2e-vitest-scenarios.yaml` - Commit and push hooks passed, including CLI tests and CLI TypeScript pre-push checks. - Documentation writer review: no docs changes needed because this only updates CI/live E2E harness behavior. --- <!-- DCO sign-off is required in this PR description, and every commit must appear as Verified in GitHub. Run: git config user.name && git config user.email --> Signed-off-by: Carlos Villela <cvillela@nvidia.com> <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **New Features** * Added an opt-in “hosted inference compatibility” mode for end-to-end onboarding and live scenarios, including hosted model/endpoint/provider routing and onboarding environment setup. * **Tests** * Updated live smoke and helper flows to reflect hosted-compatible behavior, including conditional API key checks and inference routing assertions. * Improved test consistency with a “fresh” sandbox install option. * **Chores** * Updated CI end-to-end Vitest execution to run via the hosted NVIDIA inference setup with secure API key injection and hosted model configuration. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Carlos Villela <cvillela@nvidia.com>
## Summary Restore issue NVIDIA#5800 parity package `P0-B` for merged onboard/rebuild/lifecycle bash-suite deltas only. ## Related Issues Refs NVIDIA#5800 Refs NVIDIA#5098 Refs NVIDIA#5225 Refs NVIDIA#5487 Refs NVIDIA#5410 Refs NVIDIA#5760 ## Scope gate - Package: `P0-B — Onboard/rebuild/lifecycle parity` - Included PRs all merged and touched `test/e2e`: yes — NVIDIA#5225, NVIDIA#5487, NVIDIA#5410, NVIDIA#5760 - Out of scope: unmerged/non-bash PRs; shell lane retirement / PR NVIDIA#5756 cleanup ## Parity map | ID | Source PR | Contract | Inference classification | Vitest assertion / waiver | Status | | --- | --- | --- | --- | --- | --- | | B1 | NVIDIA#5225 | Persisted sandbox-entry gateway resolution is used by lifecycle commands. | `none` | Existing `src/lib/actions/sandbox/sandbox-gateway-routing.test.ts`, `src/lib/onboard/gateway-binding.test.ts`, `src/lib/onboard/sandbox-registration.test.ts` | covered | | B2 | NVIDIA#5225 | Onboard repair/double-onboard failures include captured onboard output diagnostics. | `none` | Existing live `test/e2e-scenario/live/onboard-repair.test.ts`, `test/e2e-scenario/live/double-onboard.test.ts`; diagnostics are bash-runner-only verbosity and not a durable Vitest assertion. | waived | | B3 | NVIDIA#5487 | Plain `nemoclaw onboard` auto-detects an `in_progress` session and resumes without `--resume`; `--fresh` suppresses auto-resume. | `hosted-compatible capable` | `test/e2e-scenario/live/onboard-resume.test.ts` Phase 3.5 mutates the completed session to `in_progress`, asserts `(resume mode)` + cached skips, then asserts `--fresh` fails at injected preflight without resume banner. | covered | | B4 | NVIDIA#5410 | Rebuild resumes messaging from `messaging.plan` rather than legacy `providerCredentialHashes`; stale top-level provider hash state is not used. | `none` | `test/e2e-scenario/live/rebuild-hermes.test.ts` curated registry omits `providerCredentialHashes`; `src/lib/onboard/machine/handlers/sandbox.test.ts` refreshes registry-plan credential hashes from env on rebuild resume. | covered | | B5 | NVIDIA#5410 | Empty/staged rebuild messaging plan is preserved and token-backed channels are not rediscovered. | `none` | Existing `src/lib/actions/sandbox/rebuild-messaging-stage.test.ts` and `src/lib/onboard/machine/handlers/sandbox.test.ts`. | covered | | B6 | NVIDIA#5760 | Hosted-inference/messaging rebuild and live answer assertions tolerate model whitespace around integer `42`. | `hosted-compatible capable` | Existing `test/helpers/e2e-answer-assertions.test.ts`, consumed by `agent-turn-latency`, `full-e2e`, `launchable-smoke`, and `sandbox-operations` live Vitests. | covered | | B7 | NVIDIA#5760 | Stabilized hosted inference and messaging rebuild remain validated by live full/sandbox/launchable/rebuild targets. | `hosted-compatible capable` | Existing live targets remain unchanged; local live execution blocked by Docker daemon unavailable. Selective workflow required after PR opens. | follow-up | ## Inference mode support - Default mode for touched live targets: `hosted-compatible capable` for onboard-resume/rebuild-hermes/full/sandbox/launchable answer paths; `none` for unit/process registry and gateway routing tests. - Real inference support preserved: yes for existing hosted-compatible live targets; no new inference adapter seam added here. - Modes validated in this PR: local unit/process Vitests only; live hosted-compatible validation needs GitHub runner/secrets because local Docker daemon is unavailable. - If not validated with real inference: local `NEMOCLAW_RUN_E2E_SCENARIOS=1 npx vitest run --project e2e-scenarios-live test/e2e-scenario/live/onboard-resume.test.ts test/e2e-scenario/live/rebuild-hermes.test.ts` failed at prereq Docker daemon check before scenario assertions. ## Validation - [x] `git diff --check` - [x] `npm test -- src/lib/onboard/machine/handlers/sandbox.test.ts test/helpers/e2e-answer-assertions.test.ts src/lib/actions/sandbox/sandbox-gateway-routing.test.ts src/lib/onboard/entry-options.test.ts src/lib/onboard/sandbox-registration.test.ts src/lib/actions/sandbox/rebuild-messaging-stage.test.ts` - [x] `npm run typecheck:cli` - [x] `npm run build:cli` - [x] `npm run test-size:check` - [ ] hosted-compatible selective E2E workflow: pending PR / runner dispatch ## Follow-ups / waivers - Waiver B2: bash-only diagnostic verbosity from NVIDIA#5225 is not a durable Vitest contract; existing live tests already preserve the functional repair/double-onboard behavior. - Follow-up B7: dispatch selective live Vitest scenarios on GitHub runner with Docker + hosted inference secret after PR opens. <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Tests** * Added coverage for sandbox resume flows to verify updated Telegram credentials are picked up when resuming a rebuild. * Expanded end-to-end onboarding resume scenarios to confirm implicit resume behavior, including skipped cached steps and fresh runs starting from the expected point. * Strengthened rebuild scenario checks to ensure curated registry entries no longer include legacy credential-hash data. <!-- end of auto-generated comment: release notes by coderabbit.ai -->
## Summary Restore the Kimi-specific issue NVIDIA#5800 parity work for package `P0-D`; existing recovery and scope-upgrade package rows are explicitly mapped as pre-existing coverage and revalidated context, but not changed acceptance scope in this PR. ## Related Issues Refs NVIDIA#5800 Refs NVIDIA#5098 Refs NVIDIA#5342 Refs NVIDIA#5401 Refs NVIDIA#5406 Refs NVIDIA#5412 Refs NVIDIA#5413 Refs NVIDIA#5625 Refs NVIDIA#5760 ## Scope gate - Package: `P0-D — Recovery, Kimi, and scope-upgrade parity` - Included PRs all merged and touched `test/e2e`: yes - Changed acceptance scope in this PR: Kimi public-NVIDIA/mock parity (`D2`, `D3`) - Existing package rows revalidated without diff changes: recovery (`D1`) and scope-upgrade (`D4`) - Out of scope: unmerged/non-bash PRs; shell lane retirement / PR NVIDIA#5756 cleanup ## Parity map | ID | Source PR | Contract | Inference classification | Vitest assertion / waiver | Status | | --- | --- | --- | --- | --- | --- | | D1 | NVIDIA#5342, NVIDIA#5401 | Recovery proxy env sourcing, missing proxy-env warning, guard retention, ciao/networkInterfaces preload, and crash-loop stability are pre-existing package coverage. | `hermetic-default` | Existing `test/e2e-scenario/live/issue-2478-crash-loop-recovery.test.ts`, `test/e2e-scenario/support-tests/e2e-recovery-helpers.test.ts`; selective run `28186561267` job `issue-2478-crash-loop-recovery-vitest` passed. No diff changes here. | existing / revalidated context | | D2 | NVIDIA#5401 | Kimi remains a public-NVIDIA model/provider contract when run in trusted selective CI, while retaining mock fallback for local/untrusted validation. | `public-nvidia required` | `.github/workflows/e2e-vitest-scenarios.yaml`, `test/e2e-scenario/live/kimi-inference-compat.test.ts`, `test/e2e-scenario/support-tests/kimi-inference-compat-helpers.test.ts`, `test/e2e-script-workflow.test.ts` | covered / changed | | D3 | NVIDIA#5413, NVIDIA#5625 | Kimi multiturn tool calls split `hostname; date; uptime`, preserve tool-result flow, reject abandoned/continue traces, and normalize final punctuation. | `public-nvidia required` with mock fallback | `test/e2e-scenario/live/kimi-inference-compat-helpers.ts` trajectory assertions; selective run `28190216767` job `kimi-inference-compat-vitest` passed on the previous head; latest run `28193896380` passed on `f36fef6da`. | covered / changed | | D4 | NVIDIA#5406, NVIDIA#5412, NVIDIA#5760 | Scope-upgrade approval tolerates preapproved / not-reproduced states, denies `operator.admin` leakage, stays on gateway/no embedded fallback, and accepts whitespace-normalized `42`; this is pre-existing package coverage. | `hosted-compatible capable` | Existing `test/e2e-scenario/live/issue-4462-scope-upgrade-approval.test.ts`; selective run `28186561267` job `issue-4462-scope-upgrade-approval-vitest` passed. No diff changes here. | existing / revalidated context | ## Inference mode support - Default mode for touched live target: Kimi `mock` unless workflow selects `public-nvidia`. - Real inference support preserved: yes for Kimi public NVIDIA; yes for existing scope-upgrade hosted-compatible; not required for recovery. - Modes validated in this PR: Kimi public NVIDIA via selective workflows `28188683830`, `28190216767`; latest follow-up validation `28193896380` is running for head `f36fef6da`. Kimi helper/mock behavior via local support tests. - Source-of-truth contract: `NEMOCLAW_E2E_INFERENCE_MODE` is the canonical selector; absent selector defaults to mock for local/untrusted validation; unknown explicit values now fail closed; legacy `NEMOCLAW_KIMI_USE_MOCK=0` remains only as a temporary shell-lane compatibility alias until shell retirement. - Secret boundary: public Kimi workflow passes only `NVIDIA_API_KEY`; helper probe envs are secret-free by default; raw public NVIDIA key handoff is limited to onboard; sandbox `openclaw agent` now runs with a secret-free env and uses the configured `nvidia-prod` route. ## Validation - [x] `npx vitest run --project e2e-vitest-support test/e2e-scenario/support-tests/kimi-inference-compat-helpers.test.ts` - [x] `npx vitest run test/e2e-script-workflow.test.ts` - [x] `npm run typecheck:cli` - [x] `npm run test-conditionals:scan -- --top 25` - [x] `npx prek run --all-files --stage pre-push --skip tsc-plugin --skip tsc-js --skip tsc-cli --skip version-tag-sync --skip test-cli --skip test-plugin --skip source-shape-test-budget --skip test-file-size-budget --skip test-skills-yaml` - [x] `git diff --check` - [x] Kimi selective E2E / Vitest Scenarios on previous head: https://github.com/NVIDIA/NemoClaw/actions/runs/28190216767 - [x] Kimi selective E2E / Vitest Scenarios after review-gap fixes: https://github.com/NVIDIA/NemoClaw/actions/runs/28193896380 - [x] Existing recovery/scope rows revalidated in selective run: https://github.com/NVIDIA/NemoClaw/actions/runs/28186561267 (`issue-2478-crash-loop-recovery-vitest` ✅, `issue-4462-scope-upgrade-approval-vitest` ✅; Kimi in that stale run was superseded) - [ ] Local live mock Kimi: attempted but blocked by local Docker daemon unavailable (`Cannot connect to the Docker daemon at unix:///Users/jyaunches/.docker/run/docker.sock`). CI selective run is the live validation path for this head. ## Follow-ups / waivers - None. <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **New Features** * Added support for running Kimi compatibility e2e checks in either mock or public NVIDIA mode. * The live scenario now adapts its setup, redaction, and traffic validation based on the selected mode. * **Bug Fixes** * Improved handling of API key propagation so public NVIDIA runs use the expected credentials without exposing secrets in other paths. * **Tests** * Added coverage for mode selection, API key validation, workflow environment wiring, and the new public NVIDIA Vitest lane. <!-- end of auto-generated comment: release notes by coderabbit.ai -->
Summary
Stabilizes the failing nightly hosted-inference E2E checks by making live-answer assertions tolerant of harmless model-inserted whitespace while keeping prompt-echo guards. Fixes the channels remove + rebuild regression by preserving explicit empty messaging plans through rebuild staging so token-backed channels are not rediscovered from the environment.
Related Issue
Fixes #5759
Changes
MessagingWorkflowPlanner.buildRebuildPlanFromSandboxEntry()and stage them intoNEMOCLAW_MESSAGING_PLAN_B64during rebuild.42and reply-token checks, includingfull-e2e,sandbox-operations,launchable-smoke,agent-turn-latency,issue-4462, GPU double-onboard, and OpenClaw TUI chat-correlation paths.Type of Change
Verification
Verifiedin GitHubnpx prek run --from-ref main --to-ref HEADpassesnpm testpasses (broad runtime changes only)npm run docsbuilds without warnings (doc changes only)Verification details:
npm run build:clinpx vitest run --project cli test/helpers/e2e-answer-assertions.test.ts test/openclaw-tui-chat-correlation.test.ts src/lib/messaging/compiler/workflow-planner.test.ts src/lib/actions/sandbox/rebuild-messaging-stage.test.tsnpm run typecheck:clinpx prek run shellcheck --files test/e2e/lib/openclaw-json.sh test/e2e/test-full-e2e.sh test/e2e/test-sandbox-operations.sh test/e2e/test-launchable-smoke.sh test/e2e/test-agent-turn-latency-e2e.sh test/e2e/test-issue-4462-scope-upgrade-approval.shSigned-off-by: Carlos Villela cvillela@nvidia.com
Summary by CodeRabbit
Bug Fixes
42validations more robust to harmless whitespace/newlines in streamed output.Tests