Conversation
OpenClaw marks completed tool turns replayInvalid when side effects make replay unsafe. Treating that flag as failure changed native exit zero to one on Spark. Require an incomplete, abandoned, or timeout marker before including replay safety in diagnostics. Regression tests failed before the fix. All 65 focused tests now pass, including exit preservation. Signed-off-by: Julie Yaunches <jyaunches@nvidia.com>
|
Auto-sync is disabled for draft pull requests in this repository. Workflows must be run manually. Contributors can view more details about this message here. |
|
Navigate logical layers of code changes, visualize relationships, and explore their blast radius. No actionable comments were generated in the recent review. 🎉 ℹ️ Recent review info⚙️ Run configurationConfiguration used: Repository: NVIDIA/NemoClaw/.coderabbit.yaml Review profile: CHILL Plan: Enterprise Run ID: 📒 Files selected for processing (2)
Included review availability: This review used your included allowance. Your plan provides up to 12 included reviews per hour; 10 remain after this review. 📝 WalkthroughWalkthroughThe change updates incomplete-turn detection so ChangesTurn handling
Priority: ⬇️ Low Estimated code review effort: 2 (Simple) | ~10 minutes Change: Bug fix Suggested reviewers: Merge Risk: ⚪ Minimal · up to Completed non-replayable turns retain their output and native exit status, while timeout and incomplete turns remain rejected. No actionable merge-blocking risk was identified. 🚥 Pre-merge checks | ✅ 4 | ❌ 1❌ Failed checks (1 warning)
✅ Passed checks (4 passed)
✨ Finishing Touches 💡 2📝 Generate docstrings 💡
🛠️ Fix failing CI checks 💡
🧪 Generate unit tests (beta)
Comment |
Code Coverage OverviewLanguages: TypeScript TypeScript / code-coverage/pluginThe overall line coverage in commit 36b8c5e in the Show a line coverage summary of the most impacted files.
TypeScript / code-coverage/cliThe overall line coverage in commit 36b8c5e in the Show a line coverage summary of the most impacted files.
Updated |
Two E2E support tests still classified non-replayable turns as incomplete. Remove their obsolete success exception and rely on the shared turn classifier. Keep explicit timeout and incomplete-turn rejection, including after tool success. Reproduced both failures before repair. All 117 focused classifier, passthrough, and E2E support tests pass after repair. Signed-off-by: Julie Yaunches <jyaunches@nvidia.com>
|
The follow-up head is The new managed-image failures are also inherited. The Pi and staging-QA builders fail resolving Existing #12517 addresses the OpenSSL pins and trusted audit inputs; #12507 supplies the Undici remediation. This PR remains dependent on that existing chain. No audit or build protection has been bypassed, and local Advisor clearance is not claimed. |
Outcome
Successful OpenClaw tool turns retain their native exit code when
replayInvalid=true. Previously, NemoClaw changed a successful exit to 1 and reported an incomplete turn.Reason
OpenClaw can mark a completed turn unsafe to replay after tool side effects. Replay safety alone does not establish failure. This surfaced during Spark testing for #12502; the classifier also exists on main.
Changes
Verification
Follow-up repair: all 117 focused classifier, passthrough, and E2E support tests pass. Removed the obsolete E2E helper workaround; timeout and incomplete-turn rejection remain covered.
Focused Vitest run for
agent-json-provenance.test.tsandpassthrough-json.test.ts: three regression cases failed before the fix; all 65 tests passed afterward.npm run build:cli: passed on macOS and Sparky.Normal commit and publication hooks: passed, including secret scanning. The diff contains no secrets, API keys, or credentials.
Manual Sparky retest at
462f7fda94579dc0fa5cba8b43c3975993867422: completedwrite,read, andexec; returnedstatus=ok,replayInvalid=true, and CLI exit 0. Independently verified the created file and four-word count.The live receipt identified local
nvidia/Qwen3.6-35B-A3B-NVFP4, withrerouted=false. On the identical fresh response, the old classifier returnedreplayInvalid=trueas a failure marker; the fixed classifier returned no incomplete-turn signal.Searched production
srcconsumers ofreplayInvalid; this classifier was the only occurrence.Review notes
Local Advisor attempted on the repaired tree using the trusted base
41b9d9fa90b04281af8b98238aea4e7fc745aa8din the Lima Docker controller. It failed before any specialist review because the sandbox image pull exhausted the VM disk. Task-owned resources were removed. Full-diff self-review and focused tests completed under the authorized alternative path; hosted checks and selected E2E remain pending. This is not independent Advisor clearance.The Undici npm-audit failure is inherited: reproduced against current main lockfiles in Linux with Node 22.23.2 and npm 12.0.2. Existing dependency repair #12507 requires trusted-base prerequisite #12517. No dependency changes are mixed into this classifier PR.
The manual retest used the existing
pr12502-64gbsandbox and local inference endpoint on Sparky, a 128 GB physical host running a constrained-memory test configuration. This verifies result classification, not physical 64 GB hardware qualification. An initial attempt against the retired port 12503 failed before any tool call; the successful run used active port 12513.Signed-off-by: Julie Yaunches jyaunches@nvidia.com
Summary by CodeRabbit