Skip to content

test(e2e): tolerate Kimi final punctuation - #5625

Merged
cv merged 1 commit into
mainfrom
fix-kimi-final-text-normalization
Jun 23, 2026
Merged

cv merged 1 commit into
mainfrom
fix-kimi-final-text-normalization

Conversation

@cv

@cv cv commented Jun 23, 2026 •

Copy link
Copy Markdown
Collaborator

Summary

Relax the Kimi E2E final-answer assertion so the trajectory acceptance tolerates a missing trailing period while still requiring the expected sentence. This keeps the test focused on the tool-call split and successful final response instead of brittle punctuation from the hosted model.

Changes

  • Normalize final assistant text in test/e2e/test-kimi-inference-compat.sh before comparing the expected Kimi completion sentence.
  • Apply the same normalization in the migrated Vitest helper test/e2e-scenario/live/kimi-inference-compat-helpers.ts.

Type of Change

  • Code change (feature, bug fix, or refactor)
  • Code change with doc updates
  • Doc only (prose changes, no code sample modifications)
  • Doc only (includes code sample changes)

Verification

  • PR description includes the DCO sign-off declaration and every commit appears as Verified in GitHub
  • Git hooks passed during commit and push, or npx prek run --from-ref main --to-ref HEAD passes
  • Targeted tests pass for changed behavior
  • Full npm test passes (broad runtime changes only)
  • Tests added or updated for new or changed behavior
  • No secrets, API keys, or credentials committed
  • Docs updated for user-facing behavior changes
  • npm run docs builds without warnings (doc changes only)
  • Doc pages follow the style guide (doc changes only)
  • New doc pages include SPDX header and frontmatter (new pages only)

Signed-off-by: Carlos Villela cvillela@nvidia.com

Summary by CodeRabbit

  • Tests
    • Improved tolerance in AI inference compatibility test assertions by normalizing whitespace and handling optional trailing punctuation when comparing final assistant text outputs.

@cv cv self-assigned this Jun 23, 2026
@coderabbitai

coderabbitai Bot commented Jun 23, 2026 •

Copy link
Copy Markdown
Contributor

Review Change Stack

No actionable comments were generated in the recent review. 🎉

ℹ️ Recent review info
⚙️ Run configuration

Configuration used: Path: .coderabbit.yaml

Review profile: CHILL

Plan: Enterprise

Run ID: f58cea77-c06c-4d98-9982-35e2faac00f4

📥 Commits

Reviewing files that changed from the base of the PR and between d8fc114 and d421ad6.

📒 Files selected for processing (2)
  • test/e2e-scenario/live/kimi-inference-compat-helpers.ts
  • test/e2e/test-kimi-inference-compat.sh

📝 Walkthrough

Walkthrough

Two e2e test files for Kimi inference compatibility are updated to tolerate a missing trailing period in the final assistant text. Both the shell script (run_agent_prompt and check_trajectory_acceptance) and the TypeScript helper (assertTrajectory) now normalize the final text by trimming whitespace and stripping a trailing . before comparison.

Changes

Kimi Inference Compat — Final Text Assertion Normalization

Layer / File(s) Summary
Trailing-period normalization in final-text assertions
test/e2e/test-kimi-inference-compat.sh, test/e2e-scenario/live/kimi-inference-compat-helpers.ts
Shell run_agent_prompt uses ${final_text%.} to strip a trailing period before comparing against the expected string. check_trajectory_acceptance's embedded Python gains a normalize_final_text() helper (trim + strip trailing .), and the TypeScript helper's embedded Python applies the same normalization to the last assistantTexts entry.

Estimated code review effort

🎯 1 (Trivial) | ⏱️ ~3 minutes

Poem

A rabbit checked the text one day,
"That period at the end won't stay!"
With a trim and a snip so neat,
The test now passes—what a feat.
🐇✂️ No more dots to cause dismay!

🚥 Pre-merge checks | ✅ 4 | ❌ 1

❌ Failed checks (1 warning)

Check name Status Explanation Resolution
Docstring Coverage ⚠️ Warning Docstring coverage is 0.00% which is insufficient. The required threshold is 80.00%. Write docstrings for the functions missing them to satisfy the coverage threshold.
✅ Passed checks (4 passed)
Check name Status Explanation
Description Check ✅ Passed Check skipped - CodeRabbit’s high-level summary is enabled.
Title check ✅ Passed The title 'test(e2e): tolerate Kimi final punctuation' directly and clearly describes the main change: making E2E tests more tolerant of punctuation variations in Kimi's final responses.
Linked Issues check ✅ Passed Check skipped because no linked issues were found for this pull request.
Out of Scope Changes check ✅ Passed Check skipped because no linked issues were found for this pull request.

✏️ Tip: You can configure your own custom pre-merge checks in the settings.

✨ Finishing Touches
📝 Generate docstrings
  • Create stacked PR
  • Commit on current branch
🧪 Generate unit tests (beta)
  • Create PR with unit tests
  • Commit unit tests in branch fix-kimi-final-text-normalization

Comment @coderabbitai help to get the list of available commands.

@github-code-quality

github-code-quality Bot commented Jun 23, 2026 •

Copy link
Copy Markdown
Contributor

Code Coverage Overview

Languages: TypeScript

TypeScript / code-coverage/plugin

The overall coverage in the fix-kimi-final-text-... branch is 96%. Coverage data for the main branch is not yet available.

Show a code coverage summary of the most covered files.
File main fix-kimi-final-text-... d421ad6 +/-
nemoclaw/src/se...cret-scanner.ts — 100% —
nemoclaw/src/commands/slash.ts — 100% —
nemoclaw/src/li...bprocess-env.ts — 100% —
nemoclaw/src/bl...eprint/state.ts — 98% —
nemoclaw/src/onboard/config.ts — 98% —
nemoclaw/src/bl...int/snapshot.ts — 97% —
nemoclaw/src/bl...print/runner.ts — 95% —
nemoclaw/src/co...ration-state.ts — 94% —
nemoclaw/src/bl...ate-networks.ts — 94% —
nemoclaw/src/index.ts — 94% —

TypeScript / code-coverage/cli

The overall coverage in the fix-kimi-final-text-... branch is 46%. Coverage data for the main branch is not yet available.

Show a code coverage summary of the most covered files.
File main fix-kimi-final-text-... d421ad6 +/-
src/lib/state/o...oard-session.ts — 91% —
src/lib/inference/local.ts — 76% —
src/lib/sandbox/config.ts — 72% —
src/lib/actions...dbox/rebuild.ts — 67% —
src/lib/onboard/preflight.ts — 64% —
src/lib/actions...licy-channel.ts — 56% —
src/lib/state/sandbox.ts — 55% —
src/lib/onboard...er-gpu-patch.ts — 50% —
src/lib/policy/index.ts — 49% —
src/lib/onboard.ts — 18% —

Updated June 23, 2026 03:24 UTC
Code Coverage is in Public Preview. Learn more and provide us with your feedback.

@github-actions

Copy link
Copy Markdown
Contributor

E2E Advisor Recommendation

Required E2E: None
Optional E2E: kimi-inference-compat-vitest, kimi-inference-compat-e2e

Dispatch hint: kimi-inference-compat-vitest

Workflow run

Full advisor summary

E2E Recommendation Advisor

Base: origin/main
Head: HEAD
Confidence: high

Required E2E

  • None. No required E2E is recommended because the PR is tests-only and cannot affect installer/onboarding, sandbox lifecycle, credentials, network policy, inference routing, deployment, or real assistant runtime flows.

Optional E2E

  • kimi-inference-compat-vitest (medium): Optional validation for the touched Vitest helper used by the Kimi-compatible endpoint scenario; not merge-blocking because this PR changes test acceptance only, not runtime inference routing or assistant behavior.
  • kimi-inference-compat-e2e (medium-high): Optional validation for the touched legacy shell Kimi compatibility test in nightly E2E; useful if maintainers want to prove the shell assertion update still passes.

New E2E recommendations

  • None.

Dispatch hint

  • Workflow: E2E / Vitest Scenarios
  • jobs input: kimi-inference-compat-vitest

@github-actions

Copy link
Copy Markdown
Contributor

Vitest E2E Scenario Recommendation

Required Vitest E2E scenarios: kimi-inference-compat-vitest
Optional Vitest E2E scenarios: None

Dispatch required Vitest E2E scenarios:

  • gh workflow run e2e-vitest-scenarios.yaml --ref <pr-head-ref> --field jobs=kimi-inference-compat-vitest

Workflow run

Full Vitest E2E advisor summary

Vitest E2E Scenario Advisor

Base: origin/main
Head: HEAD
Confidence: high

Required Vitest E2E scenarios

  • kimi-inference-compat-vitest: The PR changes the shared helper used by the free-standing Kimi inference compatibility live Vitest test. That test is wired in e2e-vitest-scenarios.yaml as the kimi-inference-compat-vitest job, so dispatch that discrete job rather than the registry fan-out.
    • Dispatch: gh workflow run e2e-vitest-scenarios.yaml --ref <pr-head-ref> --field jobs=kimi-inference-compat-vitest

Optional Vitest E2E scenarios

  • None.

Relevant changed files

  • test/e2e-scenario/live/kimi-inference-compat-helpers.ts

@github-actions

github-actions Bot commented Jun 23, 2026 •

Copy link
Copy Markdown
Contributor

PR Review Advisor — No blocking findings

Merge posture: No blocking advisor findings
Primary next action: No advisor follow-up required beyond maintainer review.
Open items: 0 required · 0 warnings · 0 suggestions · 0 test follow-ups
Top item: No actionable findings

Workflow run details

This is an automated, non-binding review; it still expects maintainers and agents to respond to each required or warning item. Treat suggestions as current-PR improvements when they touch changed code; defer only with maintainer rationale or a linked follow-up. A human maintainer must make the final merge decision.

@cv
cv merged commit 96a1a8c into main Jun 23, 2026
46 checks passed
@cv
cv deleted the fix-kimi-final-text-normalization branch June 23, 2026 03:33
jyaunches added a commit that referenced this pull request Jun 25, 2026
## Summary
Restore the Kimi-specific issue #5800 parity work for package `P0-D`;
existing recovery and scope-upgrade package rows are explicitly mapped
as pre-existing coverage and revalidated context, but not changed
acceptance scope in this PR.

## Related Issues
Refs #5800
Refs #5098
Refs #5342
Refs #5401
Refs #5406
Refs #5412
Refs #5413
Refs #5625
Refs #5760

## Scope gate
- Package: `P0-D — Recovery, Kimi, and scope-upgrade parity`
- Included PRs all merged and touched `test/e2e`: yes
- Changed acceptance scope in this PR: Kimi public-NVIDIA/mock parity
(`D2`, `D3`)
- Existing package rows revalidated without diff changes: recovery
(`D1`) and scope-upgrade (`D4`)
- Out of scope: unmerged/non-bash PRs; shell lane retirement / PR #5756
cleanup

## Parity map
| ID | Source PR | Contract | Inference classification | Vitest
assertion / waiver | Status |
| --- | --- | --- | --- | --- | --- |
| D1 | #5342, #5401 | Recovery proxy env sourcing, missing proxy-env
warning, guard retention, ciao/networkInterfaces preload, and crash-loop
stability are pre-existing package coverage. | `hermetic-default` |
Existing
`test/e2e-scenario/live/issue-2478-crash-loop-recovery.test.ts`,
`test/e2e-scenario/support-tests/e2e-recovery-helpers.test.ts`;
selective run `28186561267` job `issue-2478-crash-loop-recovery-vitest`
passed. No diff changes here. | existing / revalidated context |
| D2 | #5401 | Kimi remains a public-NVIDIA model/provider contract when
run in trusted selective CI, while retaining mock fallback for
local/untrusted validation. | `public-nvidia required` |
`.github/workflows/e2e-vitest-scenarios.yaml`,
`test/e2e-scenario/live/kimi-inference-compat.test.ts`,
`test/e2e-scenario/support-tests/kimi-inference-compat-helpers.test.ts`,
`test/e2e-script-workflow.test.ts` | covered / changed |
| D3 | #5413, #5625 | Kimi multiturn tool calls split `hostname; date;
uptime`, preserve tool-result flow, reject abandoned/continue traces,
and normalize final punctuation. | `public-nvidia required` with mock
fallback | `test/e2e-scenario/live/kimi-inference-compat-helpers.ts`
trajectory assertions; selective run `28190216767` job
`kimi-inference-compat-vitest` passed on the previous head; latest run
`28193896380` passed on `f36fef6da`. | covered / changed |
| D4 | #5406, #5412, #5760 | Scope-upgrade approval tolerates
preapproved / not-reproduced states, denies `operator.admin` leakage,
stays on gateway/no embedded fallback, and accepts whitespace-normalized
`42`; this is pre-existing package coverage. | `hosted-compatible
capable` | Existing
`test/e2e-scenario/live/issue-4462-scope-upgrade-approval.test.ts`;
selective run `28186561267` job
`issue-4462-scope-upgrade-approval-vitest` passed. No diff changes here.
| existing / revalidated context |

## Inference mode support
- Default mode for touched live target: Kimi `mock` unless workflow
selects `public-nvidia`.
- Real inference support preserved: yes for Kimi public NVIDIA; yes for
existing scope-upgrade hosted-compatible; not required for recovery.
- Modes validated in this PR: Kimi public NVIDIA via selective workflows
`28188683830`, `28190216767`; latest follow-up validation `28193896380`
is running for head `f36fef6da`. Kimi helper/mock behavior via local
support tests.
- Source-of-truth contract: `NEMOCLAW_E2E_INFERENCE_MODE` is the
canonical selector; absent selector defaults to mock for local/untrusted
validation; unknown explicit values now fail closed; legacy
`NEMOCLAW_KIMI_USE_MOCK=0` remains only as a temporary shell-lane
compatibility alias until shell retirement.
- Secret boundary: public Kimi workflow passes only `NVIDIA_API_KEY`;
helper probe envs are secret-free by default; raw public NVIDIA key
handoff is limited to onboard; sandbox `openclaw agent` now runs with a
secret-free env and uses the configured `nvidia-prod` route.

## Validation
- [x] `npx vitest run --project e2e-vitest-support
test/e2e-scenario/support-tests/kimi-inference-compat-helpers.test.ts`
- [x] `npx vitest run test/e2e-script-workflow.test.ts`
- [x] `npm run typecheck:cli`
- [x] `npm run test-conditionals:scan -- --top 25`
- [x] `npx prek run --all-files --stage pre-push --skip tsc-plugin
--skip tsc-js --skip tsc-cli --skip version-tag-sync --skip test-cli
--skip test-plugin --skip source-shape-test-budget --skip
test-file-size-budget --skip test-skills-yaml`
- [x] `git diff --check`
- [x] Kimi selective E2E / Vitest Scenarios on previous head:
https://github.com/NVIDIA/NemoClaw/actions/runs/28190216767
- [x] Kimi selective E2E / Vitest Scenarios after review-gap fixes:
https://github.com/NVIDIA/NemoClaw/actions/runs/28193896380
- [x] Existing recovery/scope rows revalidated in selective run:
https://github.com/NVIDIA/NemoClaw/actions/runs/28186561267
(`issue-2478-crash-loop-recovery-vitest` ✅,
`issue-4462-scope-upgrade-approval-vitest` ✅; Kimi in that stale run was
superseded)
- [ ] Local live mock Kimi: attempted but blocked by local Docker daemon
unavailable (`Cannot connect to the Docker daemon at
unix:///Users/jyaunches/.docker/run/docker.sock`). CI selective run is
the live validation path for this head.

## Follow-ups / waivers
- None.


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **New Features**
* Added support for running Kimi compatibility e2e checks in either mock
or public NVIDIA mode.
* The live scenario now adapts its setup, redaction, and traffic
validation based on the selected mode.

* **Bug Fixes**
* Improved handling of API key propagation so public NVIDIA runs use the
expected credentials without exposing secrets in other paths.

* **Tests**
* Added coverage for mode selection, API key validation, workflow
environment wiring, and the new public NVIDIA Vitest lane.

<!-- end of auto-generated comment: release notes by coderabbit.ai -->
Hadar301 pushed a commit to Hadar301/NemoClaw-OpenShift that referenced this pull request Jul 12, 2026
## Summary
Restore the Kimi-specific issue NVIDIA#5800 parity work for package `P0-D`;
existing recovery and scope-upgrade package rows are explicitly mapped
as pre-existing coverage and revalidated context, but not changed
acceptance scope in this PR.

## Related Issues
Refs NVIDIA#5800
Refs NVIDIA#5098
Refs NVIDIA#5342
Refs NVIDIA#5401
Refs NVIDIA#5406
Refs NVIDIA#5412
Refs NVIDIA#5413
Refs NVIDIA#5625
Refs NVIDIA#5760

## Scope gate
- Package: `P0-D — Recovery, Kimi, and scope-upgrade parity`
- Included PRs all merged and touched `test/e2e`: yes
- Changed acceptance scope in this PR: Kimi public-NVIDIA/mock parity
(`D2`, `D3`)
- Existing package rows revalidated without diff changes: recovery
(`D1`) and scope-upgrade (`D4`)
- Out of scope: unmerged/non-bash PRs; shell lane retirement / PR NVIDIA#5756
cleanup

## Parity map
| ID | Source PR | Contract | Inference classification | Vitest
assertion / waiver | Status |
| --- | --- | --- | --- | --- | --- |
| D1 | NVIDIA#5342, NVIDIA#5401 | Recovery proxy env sourcing, missing proxy-env
warning, guard retention, ciao/networkInterfaces preload, and crash-loop
stability are pre-existing package coverage. | `hermetic-default` |
Existing
`test/e2e-scenario/live/issue-2478-crash-loop-recovery.test.ts`,
`test/e2e-scenario/support-tests/e2e-recovery-helpers.test.ts`;
selective run `28186561267` job `issue-2478-crash-loop-recovery-vitest`
passed. No diff changes here. | existing / revalidated context |
| D2 | NVIDIA#5401 | Kimi remains a public-NVIDIA model/provider contract when
run in trusted selective CI, while retaining mock fallback for
local/untrusted validation. | `public-nvidia required` |
`.github/workflows/e2e-vitest-scenarios.yaml`,
`test/e2e-scenario/live/kimi-inference-compat.test.ts`,
`test/e2e-scenario/support-tests/kimi-inference-compat-helpers.test.ts`,
`test/e2e-script-workflow.test.ts` | covered / changed |
| D3 | NVIDIA#5413, NVIDIA#5625 | Kimi multiturn tool calls split `hostname; date;
uptime`, preserve tool-result flow, reject abandoned/continue traces,
and normalize final punctuation. | `public-nvidia required` with mock
fallback | `test/e2e-scenario/live/kimi-inference-compat-helpers.ts`
trajectory assertions; selective run `28190216767` job
`kimi-inference-compat-vitest` passed on the previous head; latest run
`28193896380` passed on `f36fef6da`. | covered / changed |
| D4 | NVIDIA#5406, NVIDIA#5412, NVIDIA#5760 | Scope-upgrade approval tolerates
preapproved / not-reproduced states, denies `operator.admin` leakage,
stays on gateway/no embedded fallback, and accepts whitespace-normalized
`42`; this is pre-existing package coverage. | `hosted-compatible
capable` | Existing
`test/e2e-scenario/live/issue-4462-scope-upgrade-approval.test.ts`;
selective run `28186561267` job
`issue-4462-scope-upgrade-approval-vitest` passed. No diff changes here.
| existing / revalidated context |

## Inference mode support
- Default mode for touched live target: Kimi `mock` unless workflow
selects `public-nvidia`.
- Real inference support preserved: yes for Kimi public NVIDIA; yes for
existing scope-upgrade hosted-compatible; not required for recovery.
- Modes validated in this PR: Kimi public NVIDIA via selective workflows
`28188683830`, `28190216767`; latest follow-up validation `28193896380`
is running for head `f36fef6da`. Kimi helper/mock behavior via local
support tests.
- Source-of-truth contract: `NEMOCLAW_E2E_INFERENCE_MODE` is the
canonical selector; absent selector defaults to mock for local/untrusted
validation; unknown explicit values now fail closed; legacy
`NEMOCLAW_KIMI_USE_MOCK=0` remains only as a temporary shell-lane
compatibility alias until shell retirement.
- Secret boundary: public Kimi workflow passes only `NVIDIA_API_KEY`;
helper probe envs are secret-free by default; raw public NVIDIA key
handoff is limited to onboard; sandbox `openclaw agent` now runs with a
secret-free env and uses the configured `nvidia-prod` route.

## Validation
- [x] `npx vitest run --project e2e-vitest-support
test/e2e-scenario/support-tests/kimi-inference-compat-helpers.test.ts`
- [x] `npx vitest run test/e2e-script-workflow.test.ts`
- [x] `npm run typecheck:cli`
- [x] `npm run test-conditionals:scan -- --top 25`
- [x] `npx prek run --all-files --stage pre-push --skip tsc-plugin
--skip tsc-js --skip tsc-cli --skip version-tag-sync --skip test-cli
--skip test-plugin --skip source-shape-test-budget --skip
test-file-size-budget --skip test-skills-yaml`
- [x] `git diff --check`
- [x] Kimi selective E2E / Vitest Scenarios on previous head:
https://github.com/NVIDIA/NemoClaw/actions/runs/28190216767
- [x] Kimi selective E2E / Vitest Scenarios after review-gap fixes:
https://github.com/NVIDIA/NemoClaw/actions/runs/28193896380
- [x] Existing recovery/scope rows revalidated in selective run:
https://github.com/NVIDIA/NemoClaw/actions/runs/28186561267
(`issue-2478-crash-loop-recovery-vitest` ✅,
`issue-4462-scope-upgrade-approval-vitest` ✅; Kimi in that stale run was
superseded)
- [ ] Local live mock Kimi: attempted but blocked by local Docker daemon
unavailable (`Cannot connect to the Docker daemon at
unix:///Users/jyaunches/.docker/run/docker.sock`). CI selective run is
the live validation path for this head.

## Follow-ups / waivers
- None.


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **New Features**
* Added support for running Kimi compatibility e2e checks in either mock
or public NVIDIA mode.
* The live scenario now adapts its setup, redaction, and traffic
validation based on the selected mode.

* **Bug Fixes**
* Improved handling of API key propagation so public NVIDIA runs use the
expected credentials without exposing secrets in other paths.

* **Tests**
* Added coverage for mode selection, API key validation, workflow
environment wiring, and the new public NVIDIA Vitest lane.

<!-- end of auto-generated comment: release notes by coderabbit.ai -->
@wscurran wscurran added area: e2e End-to-end tests, nightly failures, or validation infrastructure chore Build, CI, dependency, or tooling maintenance labels Aug 1, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

area: e2e End-to-end tests, nightly failures, or validation infrastructure chore Build, CI, dependency, or tooling maintenance

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants