Repository navigation
test(e2e): add Deep Agents Code headless inference acceptance check (#5619) - #5789
Conversation
…VIDIA#5619) Add a skip-aware live check (07-deepagents-code-headless-inference.sh) that runs `dcode -n` inside a built Deep Agents Code sandbox and asserts: - config.toml routes through the managed https://inference.local endpoint - headless `dcode -n` returns a deterministic response or actionable provider/model error within a timeout (no hang/ambiguous failure) - no real provider/proxy credentials (nvapi-/sk-/xox.-/AKIA shapes) appear in config.toml, .env, .mcp.json, /tmp/nemoclaw-proxy-env.sh, or the captured output The script self-skips when the sandbox is not a Deep Agents Code sandbox, mirroring the existing 05/06 checks. A unit assertion in the image test registers the check and verifies its skip guard, prompt, inference route, and secret-scan content. The live green run requires a built sandbox plus the managed inference endpoint and is gated to the live e2e environment. Signed-off-by: Abhimanyu Kumar <abhimanyukumar7290@gmail.com>
ℹ️ Recent review info⚙️ Run configurationConfiguration used: Path: .coderabbit.yaml Review profile: CHILL Plan: Enterprise Run ID: 📒 Files selected for processing (2)
📝 WalkthroughWalkthroughAdds a new e2e shell check for headless ChangesDeep Agents Code headless inference validation
Estimated code review effort🎯 2 (Simple) | ⏱️ ~10 minutes Poem
🚥 Pre-merge checks | ✅ 4 | ❌ 1❌ Failed checks (1 warning)
✅ Passed checks (4 passed)
✨ Finishing Touches🧪 Generate unit tests (beta)
Comment |
There was a problem hiding this comment.
Actionable comments posted: 1
🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.
Inline comments:
In
`@test/e2e/e2e-cloud-experimental/checks/07-deepagents-code-headless-inference.sh`:
- Around line 48-54: The `sandbox_exec "cat ..."` substitutions in the
deepagents headless inference check can abort the script under `set -e` before
`fail_test` runs; make these reads non-fatal so missing files are handled by the
intended failure path. Update the `config_output` capture in this script to
tolerate a non-zero `cat` result (and apply the same pattern to the similar read
at the other referenced check) while keeping the existing `pass`/`fail_test`
logic unchanged.
🪄 Autofix (Beta)
Fix all unresolved CodeRabbit comments on this PR:
- Push a commit to this branch (recommended)
- Create a new PR with the fixes
ℹ️ Review info
⚙️ Run configuration
Configuration used: Path: .coderabbit.yaml
Review profile: CHILL
Plan: Enterprise
Run ID: 10bf3ad6-e479-414c-8595-0f10232b852e
📒 Files selected for processing (2)
test/e2e/e2e-cloud-experimental/checks/07-deepagents-code-headless-inference.shtest/langchain-deepagents-code-image.test.ts
|
✨ Thanks for adding the skip-aware headless inference check for Deep Agents Code that routes through the managed https://inference.local endpoint. This proposes a way to validate that dcode -n returns a response or deterministic error within a timeout without exposing credentials. Related open PRs: Related open issues: |
Manual PR Review Advisor resultThis PR Review Advisor analysis was run manually via Run: https://github.com/NVIDIA/NemoClaw/actions/runs/28209801709 Recommendation:
PR Review AdvisorThe new check is useful and in-scope, but its pass predicate can accept any non-empty local failure instead of proving managed inference, and two local security hardening issues should be fixed. Required before merge
Resolve or justify before merge
In-scope improvements
Test follow-ups to resolve or justify
What looks good
|
Signed-off-by: Carlos Villela <cvillela@nvidia.com>
| return execFileSync("bash", ["-c", `source "$1"; ${snippet}`, "bash", headlessCheckPath], { | ||
| encoding: "utf8", | ||
| env: { ...process.env, ...env }, | ||
| }); |
There was a problem hiding this comment.
Actionable comments posted: 3
🧹 Nitpick comments (1)
test/e2e/e2e-cloud-experimental/checks/07-deepagents-code-headless-inference.sh (1)
35-44: 🔒 Security & Privacy | 🔵 Trivial | ⚡ Quick winKeep sanitized provenance in the leak scan.
leak_scanandcombinedflatten every file into one blob, so a failure on Line 155 gives no sanitized evidence about which artifact leaked. Prefix each scanned chunk with its source path and report only the path plus a redacted snippet/hash on match; that keeps secrets out of logs while satisfying the requirement to record actionable evidence.Also applies to: 150-156
🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the rest with a brief reason, keep changes minimal, and validate. In `@test/e2e/e2e-cloud-experimental/checks/07-deepagents-code-headless-inference.sh` around lines 35 - 44, The leak scan currently concatenates artifact contents into one blob, which hides which source file triggered a match and makes the evidence hard to trace. Update sandbox_artifact_scan_command so each scanned chunk is prefixed with its source path, and adjust leak_scan/combined to emit only the path plus a redacted snippet or hash when a match is found. Use the existing sandbox_artifact_scan_command, leak_scan, and combined flow to keep provenance without exposing secrets.
🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.
Inline comments:
In
`@test/e2e/e2e-cloud-experimental/checks/07-deepagents-code-headless-inference.sh`:
- Around line 62-67: The managed-inference check is too weak because
`references_managed_inference_route` and the surrounding `dcode -n` assertions
can pass on generic mentions of provider/model/OpenAI/API key instead of proving
the active endpoint was `https://inference.local`. Tighten the shell checks in
`checks/07-deepagents-code-headless-inference.sh` by making the route match
verify the effective endpoint actually used by the CLI/config, and update the
result handling in the `dcode -n` validation path so only `PONG` or a concrete
remote/auth/HTTP failure tied to managed inference is accepted. Keep the
stricter logic anchored around `references_managed_inference_route`,
`references_managed_placeholder_key`, and the `dcode -n` status/output checks.
In `@test/langchain-deepagents-code-image.test.ts`:
- Around line 370-373: The onboarding expectation in
cloudExperimentalChecksForOnboarding for cloud-langchain-deepagents-code is
missing the new 07 headless inference check, so update the asserted list to
include the new check alongside the existing 05 and 06 entries. Make the change
in the test that validates registration so it matches the updated checklist and
confirms the new check is included.
- Around line 404-417: The test around runHeadlessCheckHelper only verifies
references_managed_inference_route and references_managed_placeholder_key
separately, so it does not enforce the actual route && key gate. Update this
Deep Agents Code config test to include a combined positive case where both
markers are present and separate negative cases for route-only and key-only
inputs, ensuring the check for the actual managed inference config expression is
covered.
---
Nitpick comments:
In
`@test/e2e/e2e-cloud-experimental/checks/07-deepagents-code-headless-inference.sh`:
- Around line 35-44: The leak scan currently concatenates artifact contents into
one blob, which hides which source file triggered a match and makes the evidence
hard to trace. Update sandbox_artifact_scan_command so each scanned chunk is
prefixed with its source path, and adjust leak_scan/combined to emit only the
path plus a redacted snippet or hash when a match is found. Use the existing
sandbox_artifact_scan_command, leak_scan, and combined flow to keep provenance
without exposing secrets.
🪄 Autofix (Beta)
Fix all unresolved CodeRabbit comments on this PR:
- Push a commit to this branch (recommended)
- Create a new PR with the fixes
ℹ️ Review info
⚙️ Run configuration
Configuration used: Path: .coderabbit.yaml
Review profile: CHILL
Plan: Enterprise
Run ID: f7e77036-e49b-4763-835b-85e562600216
📒 Files selected for processing (2)
test/e2e/e2e-cloud-experimental/checks/07-deepagents-code-headless-inference.shtest/langchain-deepagents-code-image.test.ts
There was a problem hiding this comment.
Caution
Inline review comments failed to post. This is likely due to GitHub's internal server error or limits when posting large numbers of comments. If you are seeing this consistently it is likely a permissions issue. Please check "Moderation" -> "Code review limits" under your organization settings.
Actionable comments posted: 3
🧹 Nitpick comments (1)
test/e2e/e2e-cloud-experimental/checks/07-deepagents-code-headless-inference.sh (1)
35-44: 🔒 Security & Privacy | 🔵 Trivial | ⚡ Quick winKeep sanitized provenance in the leak scan.
leak_scanandcombinedflatten every file into one blob, so a failure on Line 155 gives no sanitized evidence about which artifact leaked. Prefix each scanned chunk with its source path and report only the path plus a redacted snippet/hash on match; that keeps secrets out of logs while satisfying the requirement to record actionable evidence.Also applies to: 150-156
🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the rest with a brief reason, keep changes minimal, and validate. In `@test/e2e/e2e-cloud-experimental/checks/07-deepagents-code-headless-inference.sh` around lines 35 - 44, The leak scan currently concatenates artifact contents into one blob, which hides which source file triggered a match and makes the evidence hard to trace. Update sandbox_artifact_scan_command so each scanned chunk is prefixed with its source path, and adjust leak_scan/combined to emit only the path plus a redacted snippet or hash when a match is found. Use the existing sandbox_artifact_scan_command, leak_scan, and combined flow to keep provenance without exposing secrets.
🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.
Inline comments:
In
`@test/e2e/e2e-cloud-experimental/checks/07-deepagents-code-headless-inference.sh`:
- Around line 62-67: The managed-inference check is too weak because
`references_managed_inference_route` and the surrounding `dcode -n` assertions
can pass on generic mentions of provider/model/OpenAI/API key instead of proving
the active endpoint was `https://inference.local`. Tighten the shell checks in
`checks/07-deepagents-code-headless-inference.sh` by making the route match
verify the effective endpoint actually used by the CLI/config, and update the
result handling in the `dcode -n` validation path so only `PONG` or a concrete
remote/auth/HTTP failure tied to managed inference is accepted. Keep the
stricter logic anchored around `references_managed_inference_route`,
`references_managed_placeholder_key`, and the `dcode -n` status/output checks.
In `@test/langchain-deepagents-code-image.test.ts`:
- Around line 370-373: The onboarding expectation in
cloudExperimentalChecksForOnboarding for cloud-langchain-deepagents-code is
missing the new 07 headless inference check, so update the asserted list to
include the new check alongside the existing 05 and 06 entries. Make the change
in the test that validates registration so it matches the updated checklist and
confirms the new check is included.
- Around line 404-417: The test around runHeadlessCheckHelper only verifies
references_managed_inference_route and references_managed_placeholder_key
separately, so it does not enforce the actual route && key gate. Update this
Deep Agents Code config test to include a combined positive case where both
markers are present and separate negative cases for route-only and key-only
inputs, ensuring the check for the actual managed inference config expression is
covered.
---
Nitpick comments:
In
`@test/e2e/e2e-cloud-experimental/checks/07-deepagents-code-headless-inference.sh`:
- Around line 35-44: The leak scan currently concatenates artifact contents into
one blob, which hides which source file triggered a match and makes the evidence
hard to trace. Update sandbox_artifact_scan_command so each scanned chunk is
prefixed with its source path, and adjust leak_scan/combined to emit only the
path plus a redacted snippet or hash when a match is found. Use the existing
sandbox_artifact_scan_command, leak_scan, and combined flow to keep provenance
without exposing secrets.
🪄 Autofix (Beta)
Fix all unresolved CodeRabbit comments on this PR:
- Push a commit to this branch (recommended)
- Create a new PR with the fixes
ℹ️ Review info
⚙️ Run configuration
Configuration used: Path: .coderabbit.yaml
Review profile: CHILL
Plan: Enterprise
Run ID: f7e77036-e49b-4763-835b-85e562600216
📒 Files selected for processing (2)
test/e2e/e2e-cloud-experimental/checks/07-deepagents-code-headless-inference.shtest/langchain-deepagents-code-image.test.ts
🛑 Comments failed to post (3)
test/e2e/e2e-cloud-experimental/checks/07-deepagents-code-headless-inference.sh (1)
62-67: 🎯 Functional Correctness | 🟠 Major | 🏗️ Heavy lift
Tighten the managed-inference proof.
This can still pass without showing that
dcode -nactually hithttps://inference.local: Line 129 only checks thatconfig.tomlmentions the managed route somewhere, and Lines 104-106 accept very broad strings likeprovider,model,OpenAI, orAPI key. A local config/CLI failure that mentions those words is enough to report “reached managed inference” on Line 144.Require stronger evidence here: validate the active endpoint rather than mere presence, and only allow non-
PONGpasses for outputs that show a concrete remote/auth/HTTP failure tied to managed inference.Also applies to: 74-76, 127-145
🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the rest with a brief reason, keep changes minimal, and validate. In `@test/e2e/e2e-cloud-experimental/checks/07-deepagents-code-headless-inference.sh` around lines 62 - 67, The managed-inference check is too weak because `references_managed_inference_route` and the surrounding `dcode -n` assertions can pass on generic mentions of provider/model/OpenAI/API key instead of proving the active endpoint was `https://inference.local`. Tighten the shell checks in `checks/07-deepagents-code-headless-inference.sh` by making the route match verify the effective endpoint actually used by the CLI/config, and update the result handling in the `dcode -n` validation path so only `PONG` or a concrete remote/auth/HTTP failure tied to managed inference is accepted. Keep the stricter logic anchored around `references_managed_inference_route`, `references_managed_placeholder_key`, and the `dcode -n` status/output checks.Source: Path instructions
test/langchain-deepagents-code-image.test.ts (2)
370-373: 🎯 Functional Correctness | 🟠 Major | ⚡ Quick win
Add the new
07check to this expected onboarding list.This assertion still expects only
05and06, so it won't validate registration of the new headless inference check described by this PR. If the checklist was updated correctly, this test now fails; if it wasn't, the PR objective is still unmet.Suggested fix
expect(cloudExperimentalChecksForOnboarding("cloud-langchain-deepagents-code")).toEqual([ "test/e2e/e2e-cloud-experimental/checks/05-deepagents-code-landlock-readonly.sh", "test/e2e/e2e-cloud-experimental/checks/06-deepagents-code-python-egress.sh", + "test/e2e/e2e-cloud-experimental/checks/07-deepagents-code-headless-inference.sh", ]);📝 Committable suggestion
‼️ IMPORTANT
Carefully review the code before committing. Ensure that it accurately replaces the highlighted code, contains no missing lines, and has no issues with indentation. Thoroughly test & benchmark the code to ensure it meets the requirements.expect(cloudExperimentalChecksForOnboarding("cloud-langchain-deepagents-code")).toEqual([ "test/e2e/e2e-cloud-experimental/checks/05-deepagents-code-landlock-readonly.sh", "test/e2e/e2e-cloud-experimental/checks/06-deepagents-code-python-egress.sh", "test/e2e/e2e-cloud-experimental/checks/07-deepagents-code-headless-inference.sh", ]);🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the rest with a brief reason, keep changes minimal, and validate. In `@test/langchain-deepagents-code-image.test.ts` around lines 370 - 373, The onboarding expectation in cloudExperimentalChecksForOnboarding for cloud-langchain-deepagents-code is missing the new 07 headless inference check, so update the asserted list to include the new check alongside the existing 05 and 06 entries. Make the change in the test that validates registration so it matches the updated checklist and confirms the new check is included.
404-417: 🎯 Functional Correctness | 🟡 Minor | ⚡ Quick win
This test doesn't prove the config gate requires both markers.
The current assertions only show that each helper matches in isolation. A regression that accepts config with just the route or just the placeholder key would still pass here. Add a combined positive case plus negative route-only/key-only cases around the actual
route && keyexpression so this test enforces the intended contract.Suggested fix
it("requires the managed inference route and placeholder key in Deep Agents Code config", () => { - expect( - runHeadlessCheckHelper( - 'printf "%s" "$CONFIG" | references_managed_inference_route && printf route', - { CONFIG: 'base_url = "https://inference.local/v1"' }, - ), - ).toBe("route"); - expect( - runHeadlessCheckHelper( - 'printf "%s" "$CONFIG" | references_managed_placeholder_key && printf key', - { CONFIG: 'api_key_env = "DEEPAGENTS_CODE_OPENAI_API_KEY"' }, - ), - ).toBe("key"); + const requiresBoth = (config: string) => + runHeadlessCheckHelper( + [ + 'if printf "%s" "$CONFIG" | references_managed_inference_route', + '&& printf "%s" "$CONFIG" | references_managed_placeholder_key; then', + ' printf both;', + 'else', + ' printf missing;', + 'fi', + ].join(" "), + { CONFIG: config }, + ); + + expect( + requiresBoth([ + 'base_url = "https://inference.local/v1"', + 'api_key_env = "DEEPAGENTS_CODE_OPENAI_API_KEY"', + ].join("\n")), + ).toBe("both"); + expect(requiresBoth('base_url = "https://inference.local/v1"')).toBe("missing"); + expect(requiresBoth('api_key_env = "DEEPAGENTS_CODE_OPENAI_API_KEY"')).toBe("missing"); });📝 Committable suggestion
‼️ IMPORTANT
Carefully review the code before committing. Ensure that it accurately replaces the highlighted code, contains no missing lines, and has no issues with indentation. Thoroughly test & benchmark the code to ensure it meets the requirements.it("requires the managed inference route and placeholder key in Deep Agents Code config", () => { const requiresBoth = (config: string) => runHeadlessCheckHelper( [ 'if printf "%s" "$CONFIG" | references_managed_inference_route', '&& printf "%s" "$CONFIG" | references_managed_placeholder_key; then', ' printf both;', 'else', ' printf missing;', 'fi', ].join(" "), { CONFIG: config }, ); expect( requiresBoth([ 'base_url = "https://inference.local/v1"', 'api_key_env = "DEEPAGENTS_CODE_OPENAI_API_KEY"', ].join("\n")), ).toBe("both"); expect(requiresBoth('base_url = "https://inference.local/v1"')).toBe("missing"); expect(requiresBoth('api_key_env = "DEEPAGENTS_CODE_OPENAI_API_KEY"')).toBe("missing"); });🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the rest with a brief reason, keep changes minimal, and validate. In `@test/langchain-deepagents-code-image.test.ts` around lines 404 - 417, The test around runHeadlessCheckHelper only verifies references_managed_inference_route and references_managed_placeholder_key separately, so it does not enforce the actual route && key gate. Update this Deep Agents Code config test to include a combined positive case where both markers are present and separate negative cases for route-only and key-only inputs, ensuring the check for the actual managed inference config expression is covered.
…VIDIA#5619) (NVIDIA#5789) Closes NVIDIA#5619 Supersedes NVIDIA#5652 Adds `test/e2e/e2e-cloud-experimental/checks/07-deepagents-code-headless-inference.sh`, a skip-aware live check (mirroring the existing `05`/`06` Deep Agents Code checks) for headless `dcode -n`: - `config.toml` routes through the managed `https://inference.local` endpoint. - Headless `dcode -n "<deterministic prompt>"` returns a response or a deterministic, actionable provider/model error within a timeout — never a hang. - No provider or proxy credentials appear in `config.toml`, `.env`, `.mcp.json`, `/tmp/nemoclaw-proxy-env.sh`, or the captured output. The check self-skips when the sandbox is not a Deep Agents Code sandbox. A unit assertion in `langchain-deepagents-code-image.test.ts` registers the check and verifies its skip guard, prompt, inference route, and secret scan. Signed-off-by: Abhimanyu Kumar <abhimanyukumar7290@gmail.com> <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Tests** * Added an end-to-end Bash check for “Deep Agents Code” headless inference in a managed sandbox. * Validates routing to the expected local inference endpoint and use of managed placeholder API key references. * Improves headless execution verification with explicit timeout handling and deterministic success response checks. * Rejects ambiguous output and local-failure style outcomes; adds classification coverage for pass/actionable errors/timeouts. * Adds secret-like value scanning across sandbox config and captured runtime/proxy artifacts. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Abhimanyu Kumar <abhimanyukumar7290@gmail.com> Signed-off-by: Carlos Villela <cvillela@nvidia.com> Co-authored-by: Carlos Villela <cvillela@nvidia.com>
Closes #5619
Supersedes #5652
Adds
test/e2e/e2e-cloud-experimental/checks/07-deepagents-code-headless-inference.sh, a skip-aware live check (mirroring the existing05/06Deep Agents Code checks) for headlessdcode -n:config.tomlroutes through the managedhttps://inference.localendpoint.dcode -n "<deterministic prompt>"returns a response or a deterministic, actionable provider/model error within a timeout — never a hang.config.toml,.env,.mcp.json,/tmp/nemoclaw-proxy-env.sh, or the captured output.The check self-skips when the sandbox is not a Deep Agents Code sandbox. A unit assertion in
langchain-deepagents-code-image.test.tsregisters the check and verifies its skip guard, prompt, inference route, and secret scan.Signed-off-by: Abhimanyu Kumar abhimanyukumar7290@gmail.com
Summary by CodeRabbit