diff --git a/.agents/skills/nemoclaw-contributor-implement-issue/SKILL.md b/.agents/skills/nemoclaw-contributor-implement-issue/SKILL.md index d49d6ca705b..937b01314bd 100644 --- a/.agents/skills/nemoclaw-contributor-implement-issue/SKILL.md +++ b/.agents/skills/nemoclaw-contributor-implement-issue/SKILL.md @@ -59,6 +59,8 @@ Read current code, tests, workflows, and active guidance before editing. Load a Prefer a neutral or negative total line delta. Possible future reuse is not enough to add a mechanism. Preserve semantic regression coverage. Use runtime or E2E evidence only when a real boundary owns the behavior. +For live E2E evidence, follow [Run Maintainer E2E](../nemoclaw-maintainer-e2e/SKILL.md) for local execution. + ## Self-review Apply the shared change, state, and security contracts to the completed diff. Remove unrelated changes and avoidable machinery. diff --git a/.agents/skills/nemoclaw-maintainer-cut-release-tag/SKILL.md b/.agents/skills/nemoclaw-maintainer-cut-release-tag/SKILL.md index 679295dc0cc..2f69e5e7867 100644 --- a/.agents/skills/nemoclaw-maintainer-cut-release-tag/SKILL.md +++ b/.agents/skills/nemoclaw-maintainer-cut-release-tag/SKILL.md @@ -9,9 +9,9 @@ user_invocable: true # Cut Release Tag -Cut one signed annotated semver tag from a generated plan. Use the release scripts for tag writes -and `nemoclaw-maintainer-e2e` for maintainer-requested workflow dispatches. Do not improvise raw tag, -push, version-bump, or other release-state GitHub writes. +Cut one signed annotated semver tag from a generated plan. Use the release scripts for tag writes. +Use [Run Maintainer E2E](../nemoclaw-maintainer-e2e/SKILL.md) for maintainer-requested workflow +dispatches. Do not improvise tag, push, version-bump, or other release-state GitHub writes. Treat these as separate states: @@ -146,14 +146,8 @@ URLs, PR state, commit ranges, review state, check state, and image identities i ### 3. Present General E2E and Ask for a Decision -Use `nemoclaw-maintainer-e2e` to find the newest completed or active full E2E run. Show these details -instead of reducing the run to one passing/failing label: - -- candidate SHA and full-run SHA; -- status and conclusion; -- workflow attempt, created, started, and last-updated timestamps, plus age at inspection; -- workflow URL and `Release qualification` job URL; and -- any failed, cancelled, skipped, queued, or still-running results. +Follow [Report the Release Context](../nemoclaw-maintainer-e2e/SKILL.md#report-the-release-context) +to inspect the newest completed or active full run. Present its release context for the candidate. Offer exactly these three choices: diff --git a/.agents/skills/nemoclaw-maintainer-cut-release-tag/references/candidate-evidence.md b/.agents/skills/nemoclaw-maintainer-cut-release-tag/references/candidate-evidence.md index 1b139af5b21..925c60cebfc 100644 --- a/.agents/skills/nemoclaw-maintainer-cut-release-tag/references/candidate-evidence.md +++ b/.agents/skills/nemoclaw-maintainer-cut-release-tag/references/candidate-evidence.md @@ -375,7 +375,7 @@ When used, validate its cleanup receipts because the Brev workspace receives cre SELECTED_LAUNCHABLE_CHECK_FILE="$EVIDENCE_DIR/selected-launchable-check.json" run_or_stop "Launchable check-run selection" jq -er ' ([.[].check_runs[]? | - select(.name == "Staging Brev Launchable" and + select(.name == "Exact staging Brev Launchable" and .status == "completed" and .conclusion == "success")] | sort_by(.completed_at) | last) as $check | if $check == null then @@ -400,7 +400,7 @@ run_or_stop "Launchable job read" gh api \ run_or_stop "Launchable job validation" jq -e --arg sha "$CANDIDATE_SHA" \ --argjson run "$LAUNCHABLE_RUN_ID" --argjson job "$LAUNCHABLE_JOB_ID" ' .id == $job and .run_id == $run and .head_sha == $sha and - .name == "Staging Brev Launchable" and + .name == "Exact staging Brev Launchable" and .status == "completed" and .conclusion == "success" ' "$LAUNCHABLE_JOB_FILE" >/dev/null LAUNCHABLE_JOB_FIELDS_FILE="$EVIDENCE_DIR/launchable-job-fields.txt" diff --git a/.agents/skills/nemoclaw-maintainer-day/MERGE-GATE.md b/.agents/skills/nemoclaw-maintainer-day/MERGE-GATE.md index c453de36ab6..39b3840df10 100644 --- a/.agents/skills/nemoclaw-maintainer-day/MERGE-GATE.md +++ b/.agents/skills/nemoclaw-maintainer-day/MERGE-GATE.md @@ -63,7 +63,7 @@ A first-time fork contributor can require an **Approve and run** decision before ### Live E2E -Live E2E is not a default PR merge gate. When a maintainer requires it, use `nemoclaw-maintainer-e2e`. Evaluate its result for this PR. +Live E2E is not a default PR merge gate. When a maintainer requires it, follow [Run Maintainer E2E](../nemoclaw-maintainer-e2e/SKILL.md), then evaluate its result for this PR. ### Contributor and approver overlap diff --git a/.agents/skills/nemoclaw-maintainer-e2e/SKILL.md b/.agents/skills/nemoclaw-maintainer-e2e/SKILL.md index 2b3a1346048..d0cd3237942 100644 --- a/.agents/skills/nemoclaw-maintainer-e2e/SKILL.md +++ b/.agents/skills/nemoclaw-maintainer-e2e/SKILL.md @@ -1,6 +1,6 @@ --- name: nemoclaw-maintainer-e2e -description: Dispatches and reports trusted GitHub Actions E2E runs. Use for focused, full, staging Launchable, manual PR, and release-decision requests. +description: Runs local live E2E or dispatches and reports trusted GitHub Actions E2E. Use for local, focused, full, staging Launchable, manual PR, and release-decision requests. --- @@ -8,24 +8,32 @@ description: Dispatches and reports trusted GitHub Actions E2E runs. Use for foc # Run Maintainer E2E -Use `.github/workflows/e2e.yaml` from trusted `main`. Do not substitute local live E2E unless the maintainer explicitly requests local execution. +## Route the Request -Push runs publish `Relevant E2E`. Only a full manual run publishes `Release qualification`. That -aggregate reports the full suite; it does not decide whether a tag can proceed. A generic E2E -request does not authorize `Staging Brev Launchable`. +| Request | Procedure | +| --- | --- | +| Run working-tree source or a selected local commit | [Local Runs](references/local-runs.md) | +| Run the latest PR commit on GitHub, including failure-triggered comparison with its exact base | [Manual PR Runs](references/manual-pr.md) | +| Run the current `main` commit on GitHub | [Main Runs](references/main-runs.md) and the Launchable boundary below | +| Inspect existing evidence for a release decision | [Report the Release Context](#report-the-release-context) | +| Classify one failed GitHub Actions job | Load `nemoclaw-maintainer-classify-ci-failure`; this skill still owns dispatch and run-level reporting. | -## Route the Request +A new GitHub candidate run tests the latest PR commit or the current `main` commit. +Manual PR runs may replay the PR base after a candidate failure, as described in their procedure. +Report arbitrary historical candidate selection as unsupported. Preserve the workflow's identity and trust checks. + +Push runs select change-relevant E2E and publish `Relevant E2E`; they do not always run the full +suite. Only a full manual run publishes `Release qualification`. That aggregate reports the full +suite; it does not decide whether a tag can proceed. A generic E2E request does not authorize +`Exact staging Brev Launchable`. -- For E2E against a pull request revision, including failure-triggered comparison with its exact - base, read and follow [Manual PR Runs](references/manual-pr.md). -- To dispatch ordinary, focused, staging Launchable, or full E2E on `main`, read and follow - [Main Runs](references/main-runs.md) and the Launchable boundary below. -- For a release decision inspection, use the section below. Do not load a dispatch reference unless the maintainer requests a new run. -- For one failed job, load `nemoclaw-maintainer-classify-ci-failure` for bounded, redacted log and optional artifact evidence. This skill still owns dispatch and run-level reporting. +Use `.github/workflows/e2e.yaml` from trusted `main` for GitHub runs. Do not substitute local live +E2E unless the maintainer explicitly requests local execution. Do not load a dispatch reference for +a release inspection unless the maintainer requests a new run. ## Staging Brev Launchable Boundary -`Staging Brev Launchable` runs only for a trusted manual dispatch against `main`. Launchable +`Exact staging Brev Launchable` runs only for a trusted manual dispatch against `main`. Launchable mode selects only that job. Full mode adds it to the default E2E selection. The trusted workflow requires repository `maintain` or `admin` permission before the job's source checkout. @@ -37,7 +45,7 @@ results before it succeeds: - hosted and sandbox inference through the preinstalled full E2E suite; and - Brev workspace deletion and confirmed absence. -`Staging Brev Launchable` reads these credentials from repository Actions secrets: +`Exact staging Brev Launchable` reads these credentials from repository Actions secrets: - `BREV_API_KEY` authenticates the trusted host-side Brev CLI for workspace operations in the organization identified by `BREV_ORG_ID`. Candidate code does not receive this API key. @@ -80,13 +88,13 @@ gh run list --repo NVIDIA/NemoClaw --workflow e2e.yaml \ --jq 'map(select(.displayTitle | startswith("E2E full main"))) | first' ``` -Inspect `Release qualification`, `Staging Brev Launchable`, and every other job that is not +Inspect `Release qualification`, `Exact staging Brev Launchable`, and every other job that is not successful: ```bash gh run view --attempt --repo NVIDIA/NemoClaw \ --json jobs --jq '[.jobs[] | - select(.name == "Release qualification" or .name == "Staging Brev Launchable" or + select(.name == "Release qualification" or .name == "Exact staging Brev Launchable" or .status != "completed" or (.conclusion != null and .conclusion != "success")) | {name,status,conclusion,startedAt,completedAt,url}]' diff --git a/.agents/skills/nemoclaw-maintainer-e2e/references/local-runs.md b/.agents/skills/nemoclaw-maintainer-e2e/references/local-runs.md new file mode 100644 index 00000000000..593488faed3 --- /dev/null +++ b/.agents/skills/nemoclaw-maintainer-e2e/references/local-runs.md @@ -0,0 +1,26 @@ + + + +# Local E2E Runs + +Choose the source to test. Use the current checkout for working-tree source. Use a +detached worktree for one commit. + +Follow [Run Live E2E Locally](../../../../test/e2e/docs/README.md#run-live-e2e-locally) +for commands, host requirements, and cleanup. Review the selected test's environment +and cleanup contracts before execution. + +A local run does not reproduce GitHub full E2E. Host requirements, credentials, +target opt-ins, or unavailable services can cause tests to skip or fail. + +## Report the Result + +Return: + +- whether the run used working-tree content or a commit; +- the test file, or state that the aggregate local command ran; +- the resolved commit when applicable; +- the command result; and +- any retained resources or required cleanup. + +Do not describe a local aggregate run as GitHub full E2E or release qualification. diff --git a/.agents/skills/nemoclaw-maintainer-e2e/references/main-runs.md b/.agents/skills/nemoclaw-maintainer-e2e/references/main-runs.md index f7cba75bf3a..0a306b56912 100644 --- a/.agents/skills/nemoclaw-maintainer-e2e/references/main-runs.md +++ b/.agents/skills/nemoclaw-maintainer-e2e/references/main-runs.md @@ -20,7 +20,7 @@ Use this procedure only when the maintainer requests a new trusted `main` dispat A generic E2E request must not authorize the Brev Launchable path. Do not infer full mode from “all” or “complete.” Ask only when the request contains conflicting mode phrases. -Ordinary mode selects the default E2E suite without `Staging Brev Launchable`. Focused mode +Ordinary mode selects the default E2E suite without `Exact staging Brev Launchable`. Focused mode selects named jobs or typed targets; set only one selector input. Launchable mode runs only that job. Full mode adds it to the default suite. @@ -42,7 +42,7 @@ Before dispatch: - remove external resources that target cleanup did not remove; and - rotate or revoke any credential that candidate code could have copied. -`Staging Brev Launchable` uses `BREV_API_KEY` and `BREV_ORG_ID` during trusted host +`Exact staging Brev Launchable` uses `BREV_API_KEY` and `BREV_ORG_ID` during trusted host preparation. It exposes `NEMOCLAW_IMAGE_DISPATCH_TOKEN` only to the trusted host script as `GH_TOKEN`. It exports `NVIDIA_INFERENCE_API_KEY` into the Brev guest for full E2E. Candidate code in that guest can read the inference key. The workflow requires repository `maintain` or `admin` @@ -65,8 +65,9 @@ git fetch --prune origin main CANDIDATE_SHA="$(git rev-parse origin/main)" ``` -A new dispatch tests the `origin/main` commit resolved here. If the caller supplied another candidate -SHA, report the difference. Do not reject it or decide the release outcome. +A new dispatch tests the `origin/main` commit resolved here. It cannot select an arbitrary historical +commit. If the caller supplied another candidate SHA, report that this path cannot test it. Do not +dispatch a different commit or decide the release outcome. ## Dispatch Once @@ -162,7 +163,7 @@ Require the selected run to report `head_sha` equal to `CANDIDATE_SHA` and `stat `completed`. A successful ordinary, focused, Launchable, or full run has workflow conclusion `success`. Otherwise, return each failed, cancelled, skipped, or running job and URL. -For Launchable mode, also require one completed, successful `Staging Brev Launchable` job. +For Launchable mode, also require one completed, successful `Exact staging Brev Launchable` job. Preserve links to `launchable-e2e.json`, `full-e2e.log`, and `cleanup.json` for diagnosis. For full mode, also require one completed, successful `Release qualification` job. A skipped, cancelled, queued, or failed aggregate is not a passing full run. diff --git a/.agents/skills/nemoclaw-maintainer-e2e/references/manual-pr.md b/.agents/skills/nemoclaw-maintainer-e2e/references/manual-pr.md index fe0b9c3c2dc..f4f5a9f2b7c 100644 --- a/.agents/skills/nemoclaw-maintainer-e2e/references/manual-pr.md +++ b/.agents/skills/nemoclaw-maintainer-e2e/references/manual-pr.md @@ -4,9 +4,9 @@ # Manual PR E2E Use this mode when a maintainer requests E2E for a pull request. The trusted workflow stays on -`main` and first checks out the latest PR commit. Replay the same selector against the exact PR base -only after a candidate failure remains unresolved. The result is advisory and does not create a -required PR check. +`main` and first checks out the latest commit of an open PR. It rejects an arbitrary or historical +commit SHA. Replay the same selector against the exact PR base only after a candidate failure remains +unresolved. The result is advisory and does not create a required PR check. Manual PR E2E accepts only source branches in `NVIDIA/NemoClaw`, including for base replay. Review and adopt fork contributions onto a repository branch before dispatch. @@ -34,7 +34,7 @@ Before dispatch, review the complete candidate diff. After a failure: - remove resources that cleanup left behind; and - rotate or revoke exposed credentials when necessary. -`Staging Brev Launchable` is available only when the source is a branch in +`Exact staging Brev Launchable` is available only when the source is a branch in `NVIDIA/NemoClaw`. Its trusted host receives the Brev API key and image-dispatch token. The guest receives the NVIDIA inference API key. The protected managed-image and native-runtime qualification jobs define narrower trusted-host boundaries in the workflow. diff --git a/.agents/skills/nemoclaw-maintainer-fix-e2e-failures/SKILL.md b/.agents/skills/nemoclaw-maintainer-fix-e2e-failures/SKILL.md index 2cd8e08ba4a..36ce1928efc 100644 --- a/.agents/skills/nemoclaw-maintainer-fix-e2e-failures/SKILL.md +++ b/.agents/skills/nemoclaw-maintainer-fix-e2e-failures/SKILL.md @@ -85,7 +85,7 @@ Read [Review and Merge](references/review-and-merge.md) before reviewing, approv Observe automatic push runs and replacement attempts that the workflow starts. Never use `gh run rerun`, `gh workflow run .github/workflows/e2e.yaml`, or local live E2E to duplicate an automatic run. -Approving a first-time contributor's ordinary `pull_request` workflow after trust review is not a manual E2E dispatch. Environment approval for a secret-bearing or hardware E2E job is different: follow `nemoclaw-maintainer-e2e` only when the maintainer explicitly requests that run. +Approving a first-time contributor's ordinary `pull_request` workflow after trust review is not a manual E2E dispatch. Environment approval for a secret-bearing or hardware E2E job is different: follow [Run Maintainer E2E](../nemoclaw-maintainer-e2e/SKILL.md) only when the maintainer explicitly requests that run. Never weaken, skip, delete, relabel, or narrow coverage to make a failure disappear. Do not freeze `main`, block unrelated merges, or ask other maintainers to wait. diff --git a/.agents/skills/nemoclaw-maintainer-policies/references/release-train.md b/.agents/skills/nemoclaw-maintainer-policies/references/release-train.md index a710568ef35..1e4ebeee320 100644 --- a/.agents/skills/nemoclaw-maintainer-policies/references/release-train.md +++ b/.agents/skills/nemoclaw-maintainer-policies/references/release-train.md @@ -93,11 +93,10 @@ qualification` aggregate does not replace the candidate result. ## General E2E Decision -The general E2E decision records whether the maintainer chooses focused tests, the full suite, or the -displayed general E2E status. General E2E informs the maintainer; it does not decide whether a tag -can exist. Show the newest full run's full SHA, status, conclusion, attempt, created, started, and -last-updated timestamps, age at inspection, workflow URL, `Release qualification` URL, and any -failed, cancelled, skipped, queued, or active results. +General E2E informs the maintainer; it does not decide whether a tag can exist. Follow +[Report the Release Context](../../nemoclaw-maintainer-e2e/SKILL.md#report-the-release-context) +to inspect and report the newest full run and any maintainer-requested runs. +Record that evidence in the release decision. Offer three choices: diff --git a/.agents/skills/nemoclaw-maintainer-runtime-provider/SKILL.md b/.agents/skills/nemoclaw-maintainer-runtime-provider/SKILL.md index 4a96aed194b..eb157987f2c 100644 --- a/.agents/skills/nemoclaw-maintainer-runtime-provider/SKILL.md +++ b/.agents/skills/nemoclaw-maintainer-runtime-provider/SKILL.md @@ -80,8 +80,8 @@ Run the current type, repository, contract, activation, provider, and architectu diff affects. Local and CI checks establish contract behavior. They do not replace protected qualification or -the complete supported live E2E matrix against the commit under review. Use -`nemoclaw-maintainer-e2e` when the maintainer requests the GitHub Actions run. +the complete supported live E2E matrix against the commit under review. When a maintainer requests +a GitHub Actions run, follow [Run Maintainer E2E](../nemoclaw-maintainer-e2e/SKILL.md). ## Report diff --git a/.agents/skills/nemoclaw-maintainer-validate-launchable/SKILL.md b/.agents/skills/nemoclaw-maintainer-validate-launchable/SKILL.md index 82a28c88905..ce2904cbd49 100644 --- a/.agents/skills/nemoclaw-maintainer-validate-launchable/SKILL.md +++ b/.agents/skills/nemoclaw-maintainer-validate-launchable/SKILL.md @@ -51,7 +51,7 @@ When one or more checks ran without failure but another required check did not r ## Resolve the Candidate and Image Evidence Record the expected NemoClaw commit SHA before deployment. -Use the latest successful `Staging Brev Launchable` job for that SHA. +Use the latest successful `Exact staging Brev Launchable` job for that SHA. Record the selected workflow and job URLs and the producer run ID selected in that job's log. Download its private artifact and require `launchable-e2e.json` to report: diff --git a/.agents/skills/nemoclaw-skills-guide/SKILL.md b/.agents/skills/nemoclaw-skills-guide/SKILL.md index 5d32b6794c8..7a08c304eef 100644 --- a/.agents/skills/nemoclaw-skills-guide/SKILL.md +++ b/.agents/skills/nemoclaw-skills-guide/SKILL.md @@ -63,7 +63,7 @@ Component-specific guidance lives with the package it describes, not in a skill. | `nemoclaw-maintainer-day` | Run one daytime maintainer pass for the release version. Select a merge, salvage, security, test, conflict, or sequencing workflow. Designed for `/loop`. | | `nemoclaw-maintainer-evening` | Complete the cumulative documentation PR and release entry, show release context, and optionally start tag cutting. | | `nemoclaw-maintainer-cut-release-tag` | Verify candidate evidence, record the maintainer's E2E decision, and cut one signed semver tag. | -| `nemoclaw-maintainer-e2e` | Describe default E2E triggered by pushes to `main`, dispatch exact-revision manual PR E2E, and verify applicable workflow evidence. | +| [`nemoclaw-maintainer-e2e`](../nemoclaw-maintainer-e2e/SKILL.md) | Run exact detached commits locally, inspect automatic `main` E2E, dispatch the latest PR commit or current `main` commit, and verify applicable workflow evidence. | | `nemoclaw-maintainer-classify-ci-failure` | Classify one failed GitHub Actions job from bounded, redacted logs and an optional validated artifact. | | `nemoclaw-maintainer-analyze-ci-performance` | Analyze retained CLI test timings and base-image publication latency with bounded, read-only GitHub evidence. | | `nemoclaw-maintainer-analyze-pr-value-stream` | Measure one PR from its earliest observable branch push through merge, separate approval delay from automation time, and compare the latest revision with a target. | diff --git a/test/e2e/README.md b/test/e2e/README.md index db1021b0a7c..7d9b7f0748e 100644 --- a/test/e2e/README.md +++ b/test/e2e/README.md @@ -992,12 +992,12 @@ rm -rf -- "$evidence_dir" test ! -e "$evidence_dir" ``` -A manual run with `jobs=staging-brev-launchable` runs only `staging Brev Launchable`. +A manual run with `jobs=staging-brev-launchable` runs only `Exact staging Brev Launchable`. Push runs do not select this job. A manual trusted-`main` run with `jobs=staging-brev-launchable-identity` and an empty `targets` selector runs -only `staging Brev Launchable identity`. It builds the candidate +only `Exact staging Brev Launchable identity`. It builds the candidate image, deploys the standing Launchable, waits for workspace and SSH readiness, checks the concrete boot image and baked runtime identity, and confirms two consecutive absent workspace observations during cleanup. A passing run uploads @@ -1053,7 +1053,7 @@ without cancelling a running job. GitHub keeps at most one pending job in that group, so a newer job can replace an older pending job. For a full manual run dispatched against `main`, `Release qualification` waits -for every E2E job that does not require a separate opt-in, including `staging Brev Launchable`. The strict aggregate reports whether that full run +for every E2E job that does not require a separate opt-in, including `Exact staging Brev Launchable`. The strict aggregate reports whether that full run passed; it does not authorize or reject a tag. For a release decision, report the newest identifiable full run, its timestamps, tested commit SHA, workflow result, `Release qualification` result, and every job that did not succeed. @@ -1063,7 +1063,7 @@ rerun focused jobs, or request another full run. The release-tag skill records only the general E2E decision and any reason for proceeding with exceptional status in the signed release brief. -Separately, every release candidate requires a successful candidate `staging Brev Launchable` job. That evidence can come from a Launchable-only or +Separately, every release candidate requires a successful candidate `Exact staging Brev Launchable` job. That evidence can come from a Launchable-only or full run and cannot be waived by the maintainer's general E2E decision. The job builds the candidate image, deploys the standing Launchable, verifies the booted image and baked runtime, runs the preinstalled full E2E suite with @@ -1412,7 +1412,7 @@ A main push can queue repository-owned GPU runners or create external resources The main-run observer records attempt evidence but does not request broad failed-job reruns. Each E2E test owns any bounded operation-level retry policy. -`staging Brev Launchable` runs only for a trusted manual dispatch against `main`. +`Exact staging Brev Launchable` runs only for a trusted manual dispatch against `main`. The job reads these credentials from repository Actions secrets: - `BREV_API_KEY` authenticates the trusted host-side Brev CLI for workspace @@ -1424,13 +1424,8 @@ The job reads these credentials from repository Actions secrets: - `NVIDIA_INFERENCE_API_KEY` is exported into the Brev guest for the full E2E process. Code in the baked candidate checkout can read and use it. -`brev login` writes `BREV_API_KEY` and `BREV_ORG_ID` to -`$HOME/.brev/credentials.json` on the GitHub-hosted runner. Later trusted steps -and processes in the same job can read that file. An always-run workflow step -removes the temporary credential home and verifies its absence after the scenario. -These credentials remain valid until they expire or an administrator revokes -them in their issuing services. If cleanup fails, remove the recorded Brev -workspace. Rotate or revoke each credential to remove later access. +Follow the [Staging Brev Launchable Boundary](../../.agents/skills/nemoclaw-maintainer-e2e/SKILL.md#staging-brev-launchable-boundary) +for credential access, lifetime, and cleanup. For a same-repository PR revision, the job builds and runs the candidate commit with this same credential boundary. The PR branch must be in `NVIDIA/NemoClaw` because the image producer does not accept a sibling-repository candidate. @@ -1454,7 +1449,7 @@ Before checkout, it verifies the open PR, target repository and `main` branch, c The API must report `NVIDIA/NemoClaw` as the PR source repository. Empty `jobs` and `targets` select: -- every default-selected free-standing workflow E2E except `staging Brev Launchable`; +- every default-selected free-standing workflow E2E except `Exact staging Brev Launchable`; - every catalogue target across all credential profiles; - every shared credential-free test; and - every default registry target. diff --git a/test/e2e/docs/README.md b/test/e2e/docs/README.md index 98541331325..43eda00f452 100644 --- a/test/e2e/docs/README.md +++ b/test/e2e/docs/README.md @@ -99,7 +99,85 @@ exits 0. That exit-0 skip is specific to the typed-registry matrix; the catalogue path sets `NEMOCLAW_E2E_REQUIRE_EXECUTED_TEST=1` and exits nonzero when its selection runs no tests. -## How To Run +## Run Live E2E Locally + +Review the selected revision and local changes before running setup or live E2E on your workstation. +A detached worktree shares host privileges, credentials, and Docker access. +Run source you have not reviewed and trusted in a disposable isolated environment. +Keep workstation credentials and its Docker socket outside that environment. +Supply only test-specific credentials and follow the selected test's cleanup and revocation contract. + +Run `test:live-e2e` from the checkout whose source you want to test. The command +deletes and rebuilds `dist/` from source in that checkout before Vitest starts. +It includes tracked and untracked source inputs, runs selected test files serially, +and does not retry a failed test. It deletes direct edits under generated `dist/` +and `nemoclaw/runner-dist/` paths. + +| Goal | Checkout | Command | +| --- | --- | --- | +| Run one test file with local changes | Current working tree | `npm run test:live-e2e -- test/e2e/live/.test.ts --silent=false --reporter=default` | +| Run all locally eligible live test files | Current working tree | `npm run test:live-e2e -- --silent=false --reporter=default` | +| Run one test file at a commit | Detached worktree at the commit | Use the same focused command in that worktree. | +| Run all locally eligible live test files at a commit | Detached worktree at the commit | Use the same aggregate command in that worktree. | + +A local aggregate run is not the GitHub full E2E matrix. Tests that require another +platform, runner, credential, service, or explicit target-specific opt-in can skip +or fail locally. GitHub Actions owns those job capabilities and the strict full-run +aggregate. Interactive TUI targets require the `expect` utility on the local runner. +For trusted GitHub runs targeting the latest PR commit or current `main`, follow +[Run Maintainer E2E](../../../.agents/skills/nemoclaw-maintainer-e2e/SKILL.md). + +### Run the current working tree + +Use a repository-relative test file to select one live E2E implementation. +Add `-t` when the file contains more than one test and you need one named case: + +```bash +npm run test:live-e2e -- \ + test/e2e/live/.test.ts \ + -t '' \ + --silent=false --reporter=default +``` + +Omit `-t` to run the complete file. Omit the test file to collect every +`e2e-live` test file that the local host can run. That aggregate tests `HEAD` +only when `git status --short` is empty. Otherwise, it tests working-tree source. + +Review the selected test's environment checks and cleanup contract before you start +it. Live tests can install software and mutate Docker, OpenShell, sandbox, and +external-service state. + +### Run a commit without changing the current checkout + +Create a detached worktree, prepare that checkout, and run the selected command +inside it: + +```bash +SHA='' +COMMIT="$(git rev-parse --verify "${SHA}^{commit}")" +WORKTREE="$(mktemp -d -t nemoclaw-e2e-XXXXXXXX)" +rmdir "$WORKTREE" +git worktree add --detach "$WORKTREE" "$COMMIT" +( + cd "$WORKTREE" + npm run dev:setup + NEMOCLAW_E2E_EXPECTED_SHA="$COMMIT" npm run test:live-e2e -- \ + test/e2e/live/.test.ts \ + -t '' \ + --silent=false --reporter=default +) +``` + +Omit `-t` to run the complete file. Omit the test file for the aggregate local +run. The detached worktree selects the commit. `NEMOCLAW_E2E_EXPECTED_SHA` supplies +that identity to tests that consume it. Do not use it in a dirty checkout to claim +that a run tested only the named commit. + +The subshell returns to the primary checkout and leaves the worktree in place. +Remove external resources recorded by a failed test. Preserve any needed artifacts. +Then run `git worktree remove "$WORKTREE"`. + +## Inspect E2E Selection and Support ```bash # List canonical target ids @@ -117,16 +195,10 @@ npx vitest run --project e2e-support --silent=false --reporter=default # Validate every live test and workflow-selected integration test without running bodies npm run test:e2e-phases:check -# Opt-in live E2E targets -npm run test:live-e2e -- --silent=false --reporter=default - # Rank one or more downloaded/extracted live artifact directories npm run test:runtime-audit -- e2e-artifacts/run-1 e2e-artifacts/run-2 ``` -The aggregate local command rebuilds the CLI before Vitest starts and runs E2E -test files serially. It does not retry a failed test. - After an eligible `E2E main` push workflow completes, `E2E / Main Retry Evidence` records its conclusion and source-attempt evidence. It does not request a broad failed-job or workflow rerun. An E2E test can retry an external operation only through its checked-in bounded policy. @@ -309,7 +381,7 @@ test/e2e/ For a PR revision run, leave `jobs` and `targets` empty for all default-selected workflow E2E, catalogue profiles, shared tests, and registry targets. - `Staging Brev Launchable` requires its separate opt-in. + `Exact staging Brev Launchable` requires its separate opt-in. Keep `allow_jetson_dispatch=false` and `allow_dgx_spark_runner_queue=false` for the default selection. If the DGX Spark flag is `true`, GitHub can pause the qualification job for the `approve-dgx-spark-image-qualification` environment.