Skip to content

test(e2e): add Deep Agents Code headless inference acceptance check (#5619) - #5789

Merged
cv merged 7 commits into
NVIDIA:mainfrom
abhi-0906:feat/issue-5619-headless-check
Jun 26, 2026
Merged

cv merged 7 commits into
NVIDIA:mainfrom
abhi-0906:feat/issue-5619-headless-check

Conversation

@abhi-0906

@abhi-0906 abhi-0906 commented Jun 25, 2026 •

Copy link
Copy Markdown
Contributor

Closes #5619
Supersedes #5652

Adds test/e2e/e2e-cloud-experimental/checks/07-deepagents-code-headless-inference.sh, a skip-aware live check (mirroring the existing 05/06 Deep Agents Code checks) for headless dcode -n:

  • config.toml routes through the managed https://inference.local endpoint.
  • Headless dcode -n "<deterministic prompt>" returns a response or a deterministic, actionable provider/model error within a timeout — never a hang.
  • No provider or proxy credentials appear in config.toml, .env, .mcp.json, /tmp/nemoclaw-proxy-env.sh, or the captured output.

The check self-skips when the sandbox is not a Deep Agents Code sandbox. A unit assertion in langchain-deepagents-code-image.test.ts registers the check and verifies its skip guard, prompt, inference route, and secret scan.

Signed-off-by: Abhimanyu Kumar abhimanyukumar7290@gmail.com

Summary by CodeRabbit

  • Tests
    • Added an end-to-end Bash check for “Deep Agents Code” headless inference in a managed sandbox.
    • Validates routing to the expected local inference endpoint and use of managed placeholder API key references.
    • Improves headless execution verification with explicit timeout handling and deterministic success response checks.
    • Rejects ambiguous output and local-failure style outcomes; adds classification coverage for pass/actionable errors/timeouts.
    • Adds secret-like value scanning across sandbox config and captured runtime/proxy artifacts.

…VIDIA#5619)

Add a skip-aware live check (07-deepagents-code-headless-inference.sh) that runs
`dcode -n` inside a built Deep Agents Code sandbox and asserts:

- config.toml routes through the managed https://inference.local endpoint
- headless `dcode -n` returns a deterministic response or actionable provider/model
  error within a timeout (no hang/ambiguous failure)
- no real provider/proxy credentials (nvapi-/sk-/xox.-/AKIA shapes) appear in
  config.toml, .env, .mcp.json, /tmp/nemoclaw-proxy-env.sh, or the captured output

The script self-skips when the sandbox is not a Deep Agents Code sandbox, mirroring
the existing 05/06 checks. A unit assertion in the image test registers the check
and verifies its skip guard, prompt, inference route, and secret-scan content.

The live green run requires a built sandbox plus the managed inference endpoint and
is gated to the live e2e environment.

Signed-off-by: Abhimanyu Kumar <abhimanyukumar7290@gmail.com>
@copy-pr-bot

copy-pr-bot Bot commented Jun 25, 2026

Copy link
Copy Markdown

This pull request requires additional validation before any workflows can run on NVIDIA's runners.

Pull request vetters can view their responsibilities here.

Contributors can view more details about this message here.

@coderabbitai

coderabbitai Bot commented Jun 25, 2026 •

Copy link
Copy Markdown
Contributor

Review Change Stack

ℹ️ Recent review info
⚙️ Run configuration

Configuration used: Path: .coderabbit.yaml

Review profile: CHILL

Plan: Enterprise

Run ID: f7e77036-e49b-4763-835b-85e562600216

📥 Commits

Reviewing files that changed from the base of the PR and between e014449 and eacca57.

📒 Files selected for processing (2)
  • test/e2e/e2e-cloud-experimental/checks/07-deepagents-code-headless-inference.sh
  • test/langchain-deepagents-code-image.test.ts

📝 Walkthrough

Walkthrough

Adds a new e2e shell check for headless dcode -n execution in a Deep Agents Code sandbox, verifying managed inference routing, timeout-handled output, and secret-shaped value scanning. A Vitest case now asserts the script contains the expected headless inference markers.

Changes

Deep Agents Code headless inference validation

Layer / File(s) Summary
Headless sandbox check
test/e2e/e2e-cloud-experimental/checks/07-deepagents-code-headless-inference.sh
Adds a shell e2e check that skips unsupported sandboxes, confirms inference.local routing, runs timed headless dcode -n, scans config and captured output for secret-shaped values, and reports pass/fail counts.
Script contract test
test/langchain-deepagents-code-image.test.ts
Adds a Vitest helper for sourcing the shell script and cases for script markers, managed route and placeholder key checks, output classification, timeout validation, and secret classification.

Estimated code review effort

🎯 2 (Simple) | ⏱️ ~10 minutes

Poem

🐇 I hopped through the sandbox, quick and spry,
I watched dcode whisper, then checked the sky.
No secret carrots tucked in the logs tonight,
Just headless hops and PONG shining bright.

🚥 Pre-merge checks | ✅ 4 | ❌ 1

❌ Failed checks (1 warning)

Check name Status Explanation Resolution
Docstring Coverage ⚠️ Warning Docstring coverage is 0.00% which is insufficient. The required threshold is 80.00%. Write docstrings for the functions missing them to satisfy the coverage threshold.
✅ Passed checks (4 passed)
Check name Status Explanation
Description Check ✅ Passed Check skipped - CodeRabbit’s high-level summary is enabled.
Title check ✅ Passed The title is concise and accurately summarizes the new headless Deep Agents Code inference acceptance check.
Linked Issues check ✅ Passed The changes implement the requested headless dcode inference check, managed routing, timeout/error handling, and secret-leak scanning.
Out of Scope Changes check ✅ Passed The PR stays focused on the new e2e check and its corresponding test coverage, with no unrelated code changes.
✨ Finishing Touches
🧪 Generate unit tests (beta)
  • Create PR with unit tests

Comment @coderabbitai help to get the list of available commands.

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 1

🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

Inline comments:
In
`@test/e2e/e2e-cloud-experimental/checks/07-deepagents-code-headless-inference.sh`:
- Around line 48-54: The `sandbox_exec "cat ..."` substitutions in the
deepagents headless inference check can abort the script under `set -e` before
`fail_test` runs; make these reads non-fatal so missing files are handled by the
intended failure path. Update the `config_output` capture in this script to
tolerate a non-zero `cat` result (and apply the same pattern to the similar read
at the other referenced check) while keeping the existing `pass`/`fail_test`
logic unchanged.
🪄 Autofix (Beta)

Fix all unresolved CodeRabbit comments on this PR:

  • Push a commit to this branch (recommended)
  • Create a new PR with the fixes

ℹ️ Review info
⚙️ Run configuration

Configuration used: Path: .coderabbit.yaml

Review profile: CHILL

Plan: Enterprise

Run ID: 10bf3ad6-e479-414c-8595-0f10232b852e

📥 Commits

Reviewing files that changed from the base of the PR and between e3b8325 and 2e36216.

📒 Files selected for processing (2)
  • test/e2e/e2e-cloud-experimental/checks/07-deepagents-code-headless-inference.sh
  • test/langchain-deepagents-code-image.test.ts

Comment thread test/e2e/e2e-cloud-experimental/checks/07-deepagents-code-headless-inference.sh Outdated
@wscurran wscurran added area: e2e End-to-end tests, nightly failures, or validation infrastructure area: inference Inference routing, serving, model selection, or outputs feature PR adds or expands user-visible functionality labels Jun 25, 2026
@wscurran

Copy link
Copy Markdown
Contributor

✨ Thanks for adding the skip-aware headless inference check for Deep Agents Code that routes through the managed https://inference.local endpoint. This proposes a way to validate that dcode -n returns a response or deterministic error within a timeout without exposing credentials.


Related open PRs:


Related open issues:

@wscurran wscurran added the integration: dcode LangChain Deep Code integration behavior label Jun 25, 2026
@wscurran
wscurran requested a review from cv June 25, 2026 14:31
@cv

cv commented Jun 26, 2026

Copy link
Copy Markdown
Collaborator

Manual PR Review Advisor result

This PR Review Advisor analysis was run manually via workflow_dispatch, so the workflow did not post its usual sticky comment. Posting the advisor summary here to populate the PR with advisor feedback.

Run: https://github.com/NVIDIA/NemoClaw/actions/runs/28209801709

Recommendation: needs_rework (high confidence; findings: 3)

The new check is useful and in-scope, but its pass predicate can accept any non-empty local failure instead of proving managed inference, and two local security hardening issues should be fixed.


PR Review Advisor

The new check is useful and in-scope, but its pass predicate can accept any non-empty local failure instead of proving managed inference, and two local security hardening issues should be fixed.

Required before merge

  • Headless check can pass on any non-empty dcode failure (test/e2e/e2e-cloud-experimental/checks/07-deepagents-code-headless-inference.sh:61): The acceptance check treats any non-empty output from dcode -n as a deterministic response or actionable provider/model error. That can false-pass on local wrapper failures, parser usage text, Python tracebacks, bad timeout invocation, or other errors that occur before Deep Agents Code reaches https://inference.local/v1.
    • Impact: The linked issue requires a live test that proves headless execution reaches managed inference and produces either a successful response or deterministic provider/model error. With the current predicate, the test can go green while the managed inference path is broken or never contacted.
    • Recommendation: Replace the generic non-empty-output predicate with an explicit classifier: accept a clear successful response such as the requested PONG, or allowlisted provider/model/inference error signatures; reject local shell errors, usage output, Python tracebacks, command-not-found, and other generic non-empty failures. If practical, include evidence that the configured inference.local route was actually exercised rather than only present in config.toml.
    • Verification hint: Read lines around the headless_output handling and confirm the elif [ -n ... ] branch has been replaced with explicit success/provider-error classification.
    • Missing regression test: Add automated coverage for the classifier with representative captured outputs: PONG passes, a known provider/model error passes, and usage: dcode, Traceback, command not found, and arbitrary non-empty text fail.
    • Evidence: elif [ -n "$(printf '%s' "$headless_output" | sed 's/DCODE_EXIT:[0-9]*//' | tr -d '[:space:]')" ]; then pass "dcode -n produced a deterministic response or actionable error ..."

Resolve or justify before merge

  • Failure path can print raw Deep Agents config while scanning sensitive files (test/e2e/e2e-cloud-experimental/checks/07-deepagents-code-headless-inference.sh:51): The script intentionally reads credential-bearing locations, but the missing-route failure path prints the first 200 characters of config.toml before the leak scan result is considered. The leak regex is also narrower than the repository's canonical token families, missing examples such as nvcf-, ghp_, github_pat_, sk-proj-, sk-ant-, xapp-, and ASIA.
    • Impact: If this test catches exactly the kind of regression it is meant to detect, it may write a real credential or sensitive provider configuration prefix into E2E logs. The narrower scanner can also miss credential-shaped values in .env, .mcp.json, proxy env output, or captured runtime output.
    • Recommendation: Do not include raw config_output in failure messages; report only sanitized metadata such as byte count or that inference.local was absent. Expand the shell secret pattern to mirror the repo-wide token families used in src/lib/security/secret-patterns.ts, or otherwise document and enforce the exact provider/proxy credential families this check owns.
    • Verification hint: Read the failure branch for the inference.local check and confirm it no longer interpolates ${config_output...}; then compare SECRET_PATTERN against src/lib/security/secret-patterns.ts for the common token families.
    • Missing regression test: Extend the static Vitest contract to assert the check does not contain raw ${config_output logging and that the shell scanner covers representative ghp_, github_pat_, nvcf-, xapp-, and ASIA examples or an equivalent shared pattern source.
    • Evidence: fail_test "config.toml does not reference inference.local: ${config_output:0:200}" and SECRET_PATTERN='nvapi-[A-Za-z0-9_-]{12,}|sk-[A-Za-z0-9]{16,}|xox[baprs]-[A-Za-z0-9-]{10,}|AKIA[A-Z0-9]{16}'
  • Environment-controlled timeout is interpolated into bash -c (test/e2e/e2e-cloud-experimental/checks/07-deepagents-code-headless-inference.sh:55): DEEPAGENTS_HEADLESS_TIMEOUT flows into HEADLESS_TIMEOUT and is interpolated directly into the sandbox bash -c command as timeout ${HEADLESS_TIMEOUT} dcode ... without numeric validation.
    • Impact: A malformed or hostile environment value can inject extra shell syntax inside the sandbox command, alter the test result, or bypass the intended timeout. This is sandbox-local rather than a host escape, but the changed surface is an inference/security E2E check and should not be bypassable by shell-string injection.
    • Recommendation: Validate HEADLESS_TIMEOUT as a positive integer before constructing the command, reject invalid values, and keep the validated value quoted or passed safely. For example, accept only ^[1-9][0-9]*$ before using it with timeout.
    • Verification hint: Read the script initialization and headless_output command and confirm there is a validation guard for HEADLESS_TIMEOUT before it is interpolated into bash -c.
    • Missing regression test: Add a static or shell-level test that sets DEEPAGENTS_HEADLESS_TIMEOUT to a value containing shell metacharacters and verifies the check rejects it before running dcode; also verify a normal numeric value is accepted.
    • Evidence: HEADLESS_TIMEOUT="${DEEPAGENTS_HEADLESS_TIMEOUT:-120}" followed by timeout ${HEADLESS_TIMEOUT} dcode -n 'Reply with exactly one word: PONG'

In-scope improvements

  • None.

Test follow-ups to resolve or justify

  • Runtime validation — Classifier accepts an exact PONG response and rejects arbitrary non-empty local dcode errors. The PR is test-only, but it adds a live E2E check for a security-sensitive inference boundary. Static substring coverage is useful for registration, yet the shell check's own pass/fail and redaction predicates need behavior-specific regression coverage.
  • Runtime validation — Classifier accepts only known provider/model/inference error signatures as actionable failures. The PR is test-only, but it adds a live E2E check for a security-sensitive inference boundary. Static substring coverage is useful for registration, yet the shell check's own pass/fail and redaction predicates need behavior-specific regression coverage.
  • Runtime validation — Invalid DEEPAGENTS_HEADLESS_TIMEOUT containing shell metacharacters is rejected before sandbox execution. The PR is test-only, but it adds a live E2E check for a security-sensitive inference boundary. Static substring coverage is useful for registration, yet the shell check's own pass/fail and redaction predicates need behavior-specific regression coverage.
  • Runtime validation — Config-route failure does not print raw config.toml contents. The PR is test-only, but it adds a live E2E check for a security-sensitive inference boundary. Static substring coverage is useful for registration, yet the shell check's own pass/fail and redaction predicates need behavior-specific regression coverage.
  • Runtime validation — Credential scan catches representative nvapi, nvcf, ghp_, github_pat_, sk-proj, xapp, AKIA, and ASIA token-shaped values. The PR is test-only, but it adds a live E2E check for a security-sensitive inference boundary. Static substring coverage is useful for registration, yet the shell check's own pass/fail and redaction predicates need behavior-specific regression coverage.
  • Acceptance clause: Validate headless dcode -n inference for the supported Deep Agents Code harness. — add test evidence or identify existing coverage. The PR adds test/e2e/e2e-cloud-experimental/checks/07-deepagents-code-headless-inference.sh and a static Vitest contract, but the live predicate can pass on any non-empty output rather than proving inference.
  • Acceptance clause: Route the request through https://inference.local/v1 using the managed placeholder key. — add test evidence or identify existing coverage. The check verifies config.toml contains inference.local, and nearby agent files configure the managed route, but the script does not prove the dcode -n request actually reached that route.
  • Acceptance clause: Confirm the run produces a successful response or a deterministic, actionable model/provider error. — add test evidence or identify existing coverage. The script fails on timeout and empty output, but any non-empty local failure is accepted as deterministic/actionable.

What looks good

  • The PR is small and patches files that still exist in the current checkout.
  • The new check follows the simple skip-aware structure of the existing Deep Agents Code live checks instead of adding a new E2E framework layer.
  • The check targets important NemoClaw boundaries: managed inference routing, no hangs, and absence of provider/proxy credentials in sandbox-visible files.
  • The existing quickstart already documents dcode -n usage and managed inference routing, so no documentation drift was found.

Comment on lines +61 to +64
return execFileSync("bash", ["-c", `source "$1"; ${snippet}`, "bash", headlessCheckPath], {
encoding: "utf8",
env: { ...process.env, ...env },
});

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 3

🧹 Nitpick comments (1)
test/e2e/e2e-cloud-experimental/checks/07-deepagents-code-headless-inference.sh (1)

35-44: 🔒 Security & Privacy | 🔵 Trivial | ⚡ Quick win

Keep sanitized provenance in the leak scan.

leak_scan and combined flatten every file into one blob, so a failure on Line 155 gives no sanitized evidence about which artifact leaked. Prefix each scanned chunk with its source path and report only the path plus a redacted snippet/hash on match; that keeps secrets out of logs while satisfying the requirement to record actionable evidence.

Also applies to: 150-156

🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In
`@test/e2e/e2e-cloud-experimental/checks/07-deepagents-code-headless-inference.sh`
around lines 35 - 44, The leak scan currently concatenates artifact contents
into one blob, which hides which source file triggered a match and makes the
evidence hard to trace. Update sandbox_artifact_scan_command so each scanned
chunk is prefixed with its source path, and adjust leak_scan/combined to emit
only the path plus a redacted snippet or hash when a match is found. Use the
existing sandbox_artifact_scan_command, leak_scan, and combined flow to keep
provenance without exposing secrets.
🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

Inline comments:
In
`@test/e2e/e2e-cloud-experimental/checks/07-deepagents-code-headless-inference.sh`:
- Around line 62-67: The managed-inference check is too weak because
`references_managed_inference_route` and the surrounding `dcode -n` assertions
can pass on generic mentions of provider/model/OpenAI/API key instead of proving
the active endpoint was `https://inference.local`. Tighten the shell checks in
`checks/07-deepagents-code-headless-inference.sh` by making the route match
verify the effective endpoint actually used by the CLI/config, and update the
result handling in the `dcode -n` validation path so only `PONG` or a concrete
remote/auth/HTTP failure tied to managed inference is accepted. Keep the
stricter logic anchored around `references_managed_inference_route`,
`references_managed_placeholder_key`, and the `dcode -n` status/output checks.

In `@test/langchain-deepagents-code-image.test.ts`:
- Around line 370-373: The onboarding expectation in
cloudExperimentalChecksForOnboarding for cloud-langchain-deepagents-code is
missing the new 07 headless inference check, so update the asserted list to
include the new check alongside the existing 05 and 06 entries. Make the change
in the test that validates registration so it matches the updated checklist and
confirms the new check is included.
- Around line 404-417: The test around runHeadlessCheckHelper only verifies
references_managed_inference_route and references_managed_placeholder_key
separately, so it does not enforce the actual route && key gate. Update this
Deep Agents Code config test to include a combined positive case where both
markers are present and separate negative cases for route-only and key-only
inputs, ensuring the check for the actual managed inference config expression is
covered.

---

Nitpick comments:
In
`@test/e2e/e2e-cloud-experimental/checks/07-deepagents-code-headless-inference.sh`:
- Around line 35-44: The leak scan currently concatenates artifact contents into
one blob, which hides which source file triggered a match and makes the evidence
hard to trace. Update sandbox_artifact_scan_command so each scanned chunk is
prefixed with its source path, and adjust leak_scan/combined to emit only the
path plus a redacted snippet or hash when a match is found. Use the existing
sandbox_artifact_scan_command, leak_scan, and combined flow to keep provenance
without exposing secrets.
🪄 Autofix (Beta)

Fix all unresolved CodeRabbit comments on this PR:

  • Push a commit to this branch (recommended)
  • Create a new PR with the fixes

ℹ️ Review info
⚙️ Run configuration

Configuration used: Path: .coderabbit.yaml

Review profile: CHILL

Plan: Enterprise

Run ID: f7e77036-e49b-4763-835b-85e562600216

📥 Commits

Reviewing files that changed from the base of the PR and between e014449 and eacca57.

📒 Files selected for processing (2)
  • test/e2e/e2e-cloud-experimental/checks/07-deepagents-code-headless-inference.sh
  • test/langchain-deepagents-code-image.test.ts

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Caution

Inline review comments failed to post. This is likely due to GitHub's internal server error or limits when posting large numbers of comments. If you are seeing this consistently it is likely a permissions issue. Please check "Moderation" -> "Code review limits" under your organization settings.

Actionable comments posted: 3

🧹 Nitpick comments (1)
test/e2e/e2e-cloud-experimental/checks/07-deepagents-code-headless-inference.sh (1)

35-44: 🔒 Security & Privacy | 🔵 Trivial | ⚡ Quick win

Keep sanitized provenance in the leak scan.

leak_scan and combined flatten every file into one blob, so a failure on Line 155 gives no sanitized evidence about which artifact leaked. Prefix each scanned chunk with its source path and report only the path plus a redacted snippet/hash on match; that keeps secrets out of logs while satisfying the requirement to record actionable evidence.

Also applies to: 150-156

🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In
`@test/e2e/e2e-cloud-experimental/checks/07-deepagents-code-headless-inference.sh`
around lines 35 - 44, The leak scan currently concatenates artifact contents
into one blob, which hides which source file triggered a match and makes the
evidence hard to trace. Update sandbox_artifact_scan_command so each scanned
chunk is prefixed with its source path, and adjust leak_scan/combined to emit
only the path plus a redacted snippet or hash when a match is found. Use the
existing sandbox_artifact_scan_command, leak_scan, and combined flow to keep
provenance without exposing secrets.
🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

Inline comments:
In
`@test/e2e/e2e-cloud-experimental/checks/07-deepagents-code-headless-inference.sh`:
- Around line 62-67: The managed-inference check is too weak because
`references_managed_inference_route` and the surrounding `dcode -n` assertions
can pass on generic mentions of provider/model/OpenAI/API key instead of proving
the active endpoint was `https://inference.local`. Tighten the shell checks in
`checks/07-deepagents-code-headless-inference.sh` by making the route match
verify the effective endpoint actually used by the CLI/config, and update the
result handling in the `dcode -n` validation path so only `PONG` or a concrete
remote/auth/HTTP failure tied to managed inference is accepted. Keep the
stricter logic anchored around `references_managed_inference_route`,
`references_managed_placeholder_key`, and the `dcode -n` status/output checks.

In `@test/langchain-deepagents-code-image.test.ts`:
- Around line 370-373: The onboarding expectation in
cloudExperimentalChecksForOnboarding for cloud-langchain-deepagents-code is
missing the new 07 headless inference check, so update the asserted list to
include the new check alongside the existing 05 and 06 entries. Make the change
in the test that validates registration so it matches the updated checklist and
confirms the new check is included.
- Around line 404-417: The test around runHeadlessCheckHelper only verifies
references_managed_inference_route and references_managed_placeholder_key
separately, so it does not enforce the actual route && key gate. Update this
Deep Agents Code config test to include a combined positive case where both
markers are present and separate negative cases for route-only and key-only
inputs, ensuring the check for the actual managed inference config expression is
covered.

---

Nitpick comments:
In
`@test/e2e/e2e-cloud-experimental/checks/07-deepagents-code-headless-inference.sh`:
- Around line 35-44: The leak scan currently concatenates artifact contents into
one blob, which hides which source file triggered a match and makes the evidence
hard to trace. Update sandbox_artifact_scan_command so each scanned chunk is
prefixed with its source path, and adjust leak_scan/combined to emit only the
path plus a redacted snippet or hash when a match is found. Use the existing
sandbox_artifact_scan_command, leak_scan, and combined flow to keep provenance
without exposing secrets.
🪄 Autofix (Beta)

Fix all unresolved CodeRabbit comments on this PR:

  • Push a commit to this branch (recommended)
  • Create a new PR with the fixes

ℹ️ Review info
⚙️ Run configuration

Configuration used: Path: .coderabbit.yaml

Review profile: CHILL

Plan: Enterprise

Run ID: f7e77036-e49b-4763-835b-85e562600216

📥 Commits

Reviewing files that changed from the base of the PR and between e014449 and eacca57.

📒 Files selected for processing (2)
  • test/e2e/e2e-cloud-experimental/checks/07-deepagents-code-headless-inference.sh
  • test/langchain-deepagents-code-image.test.ts
🛑 Comments failed to post (3)
test/e2e/e2e-cloud-experimental/checks/07-deepagents-code-headless-inference.sh (1)

62-67: 🎯 Functional Correctness | 🟠 Major | 🏗️ Heavy lift

Tighten the managed-inference proof.

This can still pass without showing that dcode -n actually hit https://inference.local: Line 129 only checks that config.toml mentions the managed route somewhere, and Lines 104-106 accept very broad strings like provider, model, OpenAI, or API key. A local config/CLI failure that mentions those words is enough to report “reached managed inference” on Line 144.

Require stronger evidence here: validate the active endpoint rather than mere presence, and only allow non-PONG passes for outputs that show a concrete remote/auth/HTTP failure tied to managed inference.

Also applies to: 74-76, 127-145

🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In
`@test/e2e/e2e-cloud-experimental/checks/07-deepagents-code-headless-inference.sh`
around lines 62 - 67, The managed-inference check is too weak because
`references_managed_inference_route` and the surrounding `dcode -n` assertions
can pass on generic mentions of provider/model/OpenAI/API key instead of proving
the active endpoint was `https://inference.local`. Tighten the shell checks in
`checks/07-deepagents-code-headless-inference.sh` by making the route match
verify the effective endpoint actually used by the CLI/config, and update the
result handling in the `dcode -n` validation path so only `PONG` or a concrete
remote/auth/HTTP failure tied to managed inference is accepted. Keep the
stricter logic anchored around `references_managed_inference_route`,
`references_managed_placeholder_key`, and the `dcode -n` status/output checks.

Source: Path instructions

test/langchain-deepagents-code-image.test.ts (2)

370-373: 🎯 Functional Correctness | 🟠 Major | ⚡ Quick win

Add the new 07 check to this expected onboarding list.

This assertion still expects only 05 and 06, so it won't validate registration of the new headless inference check described by this PR. If the checklist was updated correctly, this test now fails; if it wasn't, the PR objective is still unmet.

Suggested fix
     expect(cloudExperimentalChecksForOnboarding("cloud-langchain-deepagents-code")).toEqual([
       "test/e2e/e2e-cloud-experimental/checks/05-deepagents-code-landlock-readonly.sh",
       "test/e2e/e2e-cloud-experimental/checks/06-deepagents-code-python-egress.sh",
+      "test/e2e/e2e-cloud-experimental/checks/07-deepagents-code-headless-inference.sh",
     ]);
📝 Committable suggestion

‼️ IMPORTANT
Carefully review the code before committing. Ensure that it accurately replaces the highlighted code, contains no missing lines, and has no issues with indentation. Thoroughly test & benchmark the code to ensure it meets the requirements.

    expect(cloudExperimentalChecksForOnboarding("cloud-langchain-deepagents-code")).toEqual([
      "test/e2e/e2e-cloud-experimental/checks/05-deepagents-code-landlock-readonly.sh",
      "test/e2e/e2e-cloud-experimental/checks/06-deepagents-code-python-egress.sh",
      "test/e2e/e2e-cloud-experimental/checks/07-deepagents-code-headless-inference.sh",
    ]);
🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@test/langchain-deepagents-code-image.test.ts` around lines 370 - 373, The
onboarding expectation in cloudExperimentalChecksForOnboarding for
cloud-langchain-deepagents-code is missing the new 07 headless inference check,
so update the asserted list to include the new check alongside the existing 05
and 06 entries. Make the change in the test that validates registration so it
matches the updated checklist and confirms the new check is included.

404-417: 🎯 Functional Correctness | 🟡 Minor | ⚡ Quick win

This test doesn't prove the config gate requires both markers.

The current assertions only show that each helper matches in isolation. A regression that accepts config with just the route or just the placeholder key would still pass here. Add a combined positive case plus negative route-only/key-only cases around the actual route && key expression so this test enforces the intended contract.

Suggested fix
   it("requires the managed inference route and placeholder key in Deep Agents Code config", () => {
-    expect(
-      runHeadlessCheckHelper(
-        'printf "%s" "$CONFIG" | references_managed_inference_route && printf route',
-        { CONFIG: 'base_url = "https://inference.local/v1"' },
-      ),
-    ).toBe("route");
-    expect(
-      runHeadlessCheckHelper(
-        'printf "%s" "$CONFIG" | references_managed_placeholder_key && printf key',
-        { CONFIG: 'api_key_env = "DEEPAGENTS_CODE_OPENAI_API_KEY"' },
-      ),
-    ).toBe("key");
+    const requiresBoth = (config: string) =>
+      runHeadlessCheckHelper(
+        [
+          'if printf "%s" "$CONFIG" | references_managed_inference_route',
+          '&& printf "%s" "$CONFIG" | references_managed_placeholder_key; then',
+          '  printf both;',
+          'else',
+          '  printf missing;',
+          'fi',
+        ].join(" "),
+        { CONFIG: config },
+      );
+
+    expect(
+      requiresBoth([
+        'base_url = "https://inference.local/v1"',
+        'api_key_env = "DEEPAGENTS_CODE_OPENAI_API_KEY"',
+      ].join("\n")),
+    ).toBe("both");
+    expect(requiresBoth('base_url = "https://inference.local/v1"')).toBe("missing");
+    expect(requiresBoth('api_key_env = "DEEPAGENTS_CODE_OPENAI_API_KEY"')).toBe("missing");
   });
📝 Committable suggestion

‼️ IMPORTANT
Carefully review the code before committing. Ensure that it accurately replaces the highlighted code, contains no missing lines, and has no issues with indentation. Thoroughly test & benchmark the code to ensure it meets the requirements.

  it("requires the managed inference route and placeholder key in Deep Agents Code config", () => {
    const requiresBoth = (config: string) =>
      runHeadlessCheckHelper(
        [
          'if printf "%s" "$CONFIG" | references_managed_inference_route',
          '&& printf "%s" "$CONFIG" | references_managed_placeholder_key; then',
          '  printf both;',
          'else',
          '  printf missing;',
          'fi',
        ].join(" "),
        { CONFIG: config },
      );

    expect(
      requiresBoth([
        'base_url = "https://inference.local/v1"',
        'api_key_env = "DEEPAGENTS_CODE_OPENAI_API_KEY"',
      ].join("\n")),
    ).toBe("both");
    expect(requiresBoth('base_url = "https://inference.local/v1"')).toBe("missing");
    expect(requiresBoth('api_key_env = "DEEPAGENTS_CODE_OPENAI_API_KEY"')).toBe("missing");
  });
🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@test/langchain-deepagents-code-image.test.ts` around lines 404 - 417, The
test around runHeadlessCheckHelper only verifies
references_managed_inference_route and references_managed_placeholder_key
separately, so it does not enforce the actual route && key gate. Update this
Deep Agents Code config test to include a combined positive case where both
markers are present and separate negative cases for route-only and key-only
inputs, ensuring the check for the actual managed inference config expression is
covered.

@cv
cv merged commit c119d76 into NVIDIA:main Jun 26, 2026
33 checks passed
@coderabbitai coderabbitai Bot mentioned this pull request Jun 26, 2026
8 of 21 tasks
Hadar301 pushed a commit to Hadar301/NemoClaw-OpenShift that referenced this pull request Jul 12, 2026
…VIDIA#5619) (NVIDIA#5789)

Closes NVIDIA#5619
Supersedes NVIDIA#5652

Adds
`test/e2e/e2e-cloud-experimental/checks/07-deepagents-code-headless-inference.sh`,
a skip-aware live check (mirroring the existing `05`/`06` Deep Agents
Code checks) for headless `dcode -n`:

- `config.toml` routes through the managed `https://inference.local`
endpoint.
- Headless `dcode -n "<deterministic prompt>"` returns a response or a
deterministic, actionable provider/model error within a timeout — never
a hang.
- No provider or proxy credentials appear in `config.toml`, `.env`,
`.mcp.json`, `/tmp/nemoclaw-proxy-env.sh`, or the captured output.

The check self-skips when the sandbox is not a Deep Agents Code sandbox.
A unit assertion in `langchain-deepagents-code-image.test.ts` registers
the check and verifies its skip guard, prompt, inference route, and
secret scan.

Signed-off-by: Abhimanyu Kumar <abhimanyukumar7290@gmail.com>


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **Tests**
* Added an end-to-end Bash check for “Deep Agents Code” headless
inference in a managed sandbox.
* Validates routing to the expected local inference endpoint and use of
managed placeholder API key references.
* Improves headless execution verification with explicit timeout
handling and deterministic success response checks.
* Rejects ambiguous output and local-failure style outcomes; adds
classification coverage for pass/actionable errors/timeouts.
* Adds secret-like value scanning across sandbox config and captured
runtime/proxy artifacts.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Abhimanyu Kumar <abhimanyukumar7290@gmail.com>
Signed-off-by: Carlos Villela <cvillela@nvidia.com>
Co-authored-by: Carlos Villela <cvillela@nvidia.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

area: e2e End-to-end tests, nightly failures, or validation infrastructure area: inference Inference routing, serving, model selection, or outputs feature PR adds or expands user-visible functionality integration: dcode LangChain Deep Code integration behavior

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Validate Deep Agents Code headless dcode inference

4 participants