Skip to content

fix(onboard): bound chat tool-call validation output - #4532

Merged
cv merged 2 commits into
mainfrom
fix/ollama-tool-probe-output-cap
May 29, 2026
Merged

cv merged 2 commits into
mainfrom
fix/ollama-tool-probe-output-cap

Conversation

@cv

@cv cv commented May 29, 2026 •

Copy link
Copy Markdown
Collaborator

Summary

Bounds strict chat-completions tool-call validation requests by setting max_tokens and stream: false. This keeps slow local inference models from continuing generation until the host-side curl process timeout kills onboarding validation.

Related Issue

Refs #4501

Changes

  • Add max_tokens: 64 and stream: false to the strict chat-completions tool-call probe payload in src/lib/inference/onboard-probes.ts.
  • Update the strict retry probe test to assert the bounded validation payload.

Type of Change

  • Code change (feature, bug fix, or refactor)
  • Code change with doc updates
  • Doc only (prose changes, no code sample modifications)
  • Doc only (includes code sample changes)

Verification

  • npx prek run --all-files passes
  • npm test passes
  • Tests added or updated for new or changed behavior
  • No secrets, API keys, or credentials committed
  • Docs updated for user-facing behavior changes
  • npm run docs builds without warnings (doc changes only)
  • Doc pages follow the style guide (doc changes only)
  • New doc pages include SPDX header and frontmatter (new pages only)

Signed-off-by: Carlos Villela cvillela@nvidia.com

Summary by CodeRabbit

  • Bug Fixes

    • Constrained probe validation requests to limit output and disable streaming, improving stability of onboarding checks.
  • Tests

    • Updated tests to assert specific request payload fields (required tool choice, max tokens, non-streaming) for stricter validation.

Review Change Stack

@cv cv self-assigned this May 29, 2026
@coderabbitai

coderabbitai Bot commented May 29, 2026 •

Copy link
Copy Markdown
Contributor

No actionable comments were generated in the recent review. 🎉

ℹ️ Recent review info
⚙️ Run configuration

Configuration used: Path: .coderabbit.yaml

Review profile: CHILL

Plan: Enterprise

Run ID: a696f65f-9b37-4bc4-8689-9689493254a0

📥 Commits

Reviewing files that changed from the base of the PR and between 592f0b0 and 6ced107.

📒 Files selected for processing (2)
  • src/lib/inference/onboard-probes.test.ts
  • src/lib/inference/onboard-probes.ts
🚧 Files skipped from review as they are similar to previous changes (2)
  • src/lib/inference/onboard-probes.test.ts
  • src/lib/inference/onboard-probes.ts

📝 Walkthrough

Walkthrough

The Chat Completions tool-call validation request now includes max_tokens: 256 and stream: false, and the strict retry test parses tmpDir/request-2.json to assert tool_choice: "required", max_tokens: 256, and stream: false.

Changes

Chat Completions Tool-Call Probe Output Constraints

Layer / File(s) Summary
Constrain validation request output and verify in test
src/lib/inference/onboard-probes.ts, src/lib/inference/onboard-probes.test.ts
The probe adds max_tokens: 256 and stream: false to the Chat Completions validation request, and the strict retry test is updated to parse the saved retry request JSON and assert tool_choice: "required", max_tokens: 256, and stream: false.

Estimated code review effort

🎯 3 (Moderate) | ⏱️ ~20 minutes

Possibly related PRs

  • NVIDIA/NemoClaw#4250: Both PRs touch the Ollama/onboarding Chat Completions tool-calling validation; #4250 changes when the strict tool-calling probe is enabled while this PR adjusts the probe request payload and its test.

Suggested labels

fix, enhancement: inference

Poem

🐰 I bounded the hop and shortened the stream,
Tokens trimmed neat like a midday dream.
A tiny JSON checked with care,
Tool choice set — no endless fare.
Hop on, tests pass — carrot-time gleam!

🚥 Pre-merge checks | ✅ 4 | ❌ 1

❌ Failed checks (1 warning)

Check name Status Explanation Resolution
Docstring Coverage ⚠️ Warning Docstring coverage is 0.00% which is insufficient. The required threshold is 80.00%. Write docstrings for the functions missing them to satisfy the coverage threshold.
✅ Passed checks (4 passed)
Check name Status Explanation
Description Check ✅ Passed Check skipped - CodeRabbit’s high-level summary is enabled.
Title check ✅ Passed The title 'fix(onboard): bound chat tool-call validation output' directly and clearly summarizes the main change: adding bounds (max_tokens and stream settings) to the chat tool-call validation request during onboarding.
Linked Issues check ✅ Passed Check skipped because no linked issues were found for this pull request.
Out of Scope Changes check ✅ Passed Check skipped because no linked issues were found for this pull request.

✏️ Tip: You can configure your own custom pre-merge checks in the settings.

✨ Finishing Touches
📝 Generate docstrings
  • Create stacked PR
  • Commit on current branch
🧪 Generate unit tests (beta)
  • Create PR with unit tests
  • Commit unit tests in branch fix/ollama-tool-probe-output-cap

Warning

There were issues while running some tools. Please review the errors and either fix the tool's configuration or disable the tool if it's a critical failure.

🔧 ESLint

If the error stems from missing dependencies, add them to the package.json file. For unrecoverable errors (e.g., due to private dependencies), disable the tool in the CodeRabbit configuration.

ESLint skipped: no ESLint configuration detected in root package.json. To enable, add eslint to devDependencies.


Comment @coderabbitai help to get the list of available commands and usage tips.

@github-actions

github-actions Bot commented May 29, 2026 •

Copy link
Copy Markdown
Contributor

E2E Advisor Recommendation

Required E2E: gpu-e2e
Optional E2E: onboard-inference-smoke-e2e, kimi-inference-compat-e2e

Dispatch hint: gpu-e2e

Auto-dispatched E2E: gpu-e2e via nightly-e2e.yaml at 6ced1078548e9f34d950714f0cedd24bc98f3e9f — nightly run

Workflow run

Full advisor summary

E2E Recommendation Advisor

Base: origin/main
Head: HEAD
Confidence: high

Required E2E

  • gpu-e2e (high; requires enabled NVIDIA GPU runner and pulls/runs a local Ollama model): Most direct existing real-user-flow coverage for this change: installs/onboards with NEMOCLAW_PROVIDER=ollama, which exercises strict local Chat Completions tool-calling validation before proving sandbox inference through inference.local. This validates the new bounded strict probe payload against a real local Ollama runtime.

Optional E2E

  • onboard-inference-smoke-e2e (low; hermetic Node/CLI regression without live cloud or GPU): Useful fast regression guard for onboard inference validation fail-closed behavior and diagnostics when chat/completions is broken. It does not specifically exercise strict tool-calling payloads, so it is adjacent rather than sufficient by itself.
  • kimi-inference-compat-e2e (medium; creates sandbox and runs mock-compatible endpoint flow): Optional compatibility signal for OpenAI-compatible endpoint onboarding and downstream tool-call-oriented assistant behavior with a hermetic mock, but the onboard validation path is not the exact requireChatCompletionsToolCalling path changed here.

New E2E recommendations

  • strict Chat Completions tool-call probe payload (medium): Existing E2E coverage does not provide a cheap hermetic install/onboard flow that directly asserts strict Chat Completions validation sends bounded non-streaming payloads and succeeds/fails against a mock OpenAI-compatible endpoint without requiring GPU/Ollama.
    • Suggested test: Add a PR-safe hermetic E2E that onboards a custom/local OpenAI-compatible mock requiring structured tool_calls, captures the validation request, and asserts tool_choice=required, max_tokens=256, stream=false, and bounded retry behavior.

Dispatch hint

  • Workflow: .github/workflows/nightly-e2e.yaml
  • jobs input: gpu-e2e

@github-actions

github-actions Bot commented May 29, 2026 •

Copy link
Copy Markdown
Contributor

E2E Scenario Advisor Recommendation

Required scenario E2E: None
Optional scenario E2E: None

Workflow run

Full scenario advisor summary

E2E Scenario Advisor

Base: origin/main
Head: HEAD
Confidence: high

Required scenario E2E

  • None. No scenario workflow, scenario metadata, scenario runtime, or validation-suite files changed.

Optional scenario E2E

  • None.

Relevant changed files

  • None.

@github-actions

github-actions Bot commented May 29, 2026 •

Copy link
Copy Markdown
Contributor

PR Review Advisor

Findings: 0 needs attention, 0 worth checking, 0 nice ideas
Since last review: 0 prior items resolved, 0 still apply, 0 new items found

Workflow run details

This is an automated advisory review. A human maintainer must make the final merge decision.

@github-actions

Copy link
Copy Markdown
Contributor

Selective E2E Results — ⚠️ No requested jobs ran

Run: 26649320794
Target ref: 6ced1078548e9f34d950714f0cedd24bc98f3e9f
Workflow ref: main
Requested jobs: gpu-e2e
Summary: 0 passed, 0 failed, 1 skipped

Job Result
gpu-e2e ⏭️ skipped

@jyaunches

Copy link
Copy Markdown
Contributor

/run gpu-e2e

@cjagwani cjagwani left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM

@jyaunches

Copy link
Copy Markdown
Contributor

/run kimi-inference-compat-e2e onboard-inference-smoke-e2e

@cv
cv merged commit 6a0c022 into main May 29, 2026
30 checks passed
@cv
cv deleted the fix/ollama-tool-probe-output-cap branch May 29, 2026 17:13

@jyaunches jyaunches left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Requesting changes — E2E coverage gap

⚠️ Note on stance: We are being extremely cautious here given proximity to Computex. The fix itself is sound and low-risk in isolation, but the onboarding probe path is a launch-blocker if it regresses, so we are holding the bar higher than usual on regression coverage before merge.

Summary of findings

The change is well-scoped (adds max_tokens: 256 and stream: false to probeChatCompletionsToolCalling) and the unit test correctly asserts the retry payload shape. CodeRabbit and the PR Review Advisor both report 0 actionable findings. However, two coverage concerns warrant attention before merge:

1. The E2E Advisor flagged a missing hermetic regression guard — not yet added

Quoting the advisor's own recommendation on this PR:

strict Chat Completions tool-call probe payload (medium): Existing E2E coverage does not provide a cheap hermetic install/onboard flow that directly asserts strict Chat Completions validation sends bounded non-streaming payloads and succeeds/fails against a mock OpenAI-compatible endpoint without requiring GPU/Ollama.

Suggested test: Add a PR-safe hermetic E2E that onboards a custom/local OpenAI-compatible mock requiring structured tool_calls, captures the validation request, and asserts tool_choice=required, max_tokens=256, stream=false, and bounded retry behavior.

This was not added in this PR. Verified:

  • PR diff = 2 files only (onboard-probes.ts, onboard-probes.test.ts).
  • No new entry in test/e2e/ or .github/workflows/regression-e2e.yaml.
  • No tracking issue filed.

requireChatCompletionsToolCalling: true is set in exactly one callsite (src/lib/onboard.ts:6012, the local Ollama onboard path), so gpu-e2e is currently the only test that exercises this code path end-to-end. That's a single point of failure for a probe that gates onboarding success.

2. Thinking-model carve-out gap

getChatCompletionsProbePayload already special-cases DeepSeek-V4-Pro and Kimi-K2.6 with chat_template_kwargs: { thinking: false } and tuned max_tokens (8192 / 8). The patched probeChatCompletionsToolCalling, however, applies the same 256-token cap to every model with no thinking: false hint.

Failure mode: a reasoning model that emits a thinking trace before the tool call could exhaust 256 tokens of internal reasoning and never emit tool_calls → false-negative onboarding probe failure for those models. Today this is theoretical (only Ollama hits this gate), but if the gate is ever extended to other providers it becomes a regression vector.

Requested changes before merge

Required:

  1. Add a hermetic E2E (per the advisor's suggestion) under test/e2e/ — onboard against an OpenAI-compat mock requiring structured tool_calls, assert the bounded-payload shape and retry behavior. Wire it into regression-e2e.yaml.

Acceptable alternative if (1) is too large for this PR:
2. File a follow-up issue tagged e2e-gap capturing the advisor's hermetic-test suggestion AND the thinking-model carve-out concern. Reference the issue from a comment in probeChatCompletionsToolCalling. Merge this PR only after the issue is filed and linked.

Independent of (1)/(2):
3. Add a brief code comment near the new max_tokens: 256 noting that the cap assumes non-reasoning output and may need a thinking-model carve-out if requireChatCompletionsToolCalling is ever extended beyond Ollama.

Why we're holding this line

Pre-Computex, the onboarding flow is the most user-visible failure surface we have. A bounded-output cap that's too tight, or a stream: false mismatch, would silently break onboarding for whichever models slip into the strict probe path next. The unit test catches payload-shape regressions, but not behavioral regressions against a real (or hermetic) endpoint. Closing that gap before merge — even with a follow-up issue — keeps us covered through the launch window.

@cv

cv commented May 29, 2026

Copy link
Copy Markdown
Collaborator Author

Posted follow-up issue #4537 to track the hermetic E2E gap and thinking-model carve-out concern from the review: #4537

@coderabbitai coderabbitai Bot mentioned this pull request May 29, 2026
5 of 12 tasks
@wscurran wscurran added the bug-fix PR fixes a bug or regression label Jun 8, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

bug-fix PR fixes a bug or regression

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants