Skip to content

[Refactor] Move the Linux CI jobs to OSDC runners - #4486

Draft
huydhn wants to merge 2 commits into
mainfrom
osdc/ci-runners
Draft

huydhn wants to merge 2 commits into
mainfrom
osdc/ci-runners

Conversation

@huydhn

@huydhn huydhn commented Sep 25, 2026

Copy link
Copy Markdown

Moves the Linux test, docs and lint jobs across 11 workflows onto OSDC pods.

before after
linux.g5.4xlarge.nvidia.gpu mt-l-x86aavx2-11-41-a10g
linux.g6.4xlarge.experimental.nvidia.gpu mt-l-x86aavx2-11-41-l4
linux.12xlarge mt-l-x86iavx512-48-384
linux.4xlarge mt-l-x86iavx512-16-128

Each pod sits on the same instance the label it replaces resolved to, so GPU model and count are unchanged and only the Kubernetes overhead comes off.

linux_job_v2 → v3 comes with it, per job rather than per file: OSDC rejects a job that has no container, and v3 supplies pytorch/almalinux-builder for the arch. Only jobs whose runner actually moves are converted.

The lint jobs set no runner: at all, so they were quietly taking v2's linux.2xlarge default. They now pin an OSDC pod explicitly.

Two jobs are deliberately left on EC2:

  • test-linux-libs.yml unittests-isaaclab calls vmoens/test-infra/.../isaac_linux_job_v2.yml, a fork's reusable workflow. It cannot be confirmed to set a container, and a container-less job is rejected outright on OSDC.
  • docs.yml's gh-pages upload is still a linux_job.yml (v1) caller. Moving it is part of the separate v1-removal work.

Wheel builds are in a separate PR.

Authored with Claude Code.

@pytorch-bot

pytorch-bot Bot commented Sep 25, 2026 •

Copy link
Copy Markdown

🔗 Helpful Links

🧪 See artifacts and rendered test results at hud.pytorch.org/pr/pytorch/rl/4486

Note: Links to docs will display an error until the docs builds have been completed.

❌ 4 New Failures, 16 Cancelled Jobs, 1 Unclassified Failure

As of commit 65f72b5 with merge base 1441844 (image):

NEW FAILURES - The following jobs have failed:

UNCLASSIFIED FAILURE - DrCI could not classify the following job because the workflow did not run on the merge base. The failure may be pre-existing on trunk or introduced by this PR:

  • Unit-tests on Linux / tests-optdeps-smoke (3.12, 13.0) / linux-job (gh) (this job did not run on the merge base, so DrCI cannot tell whether the failure is pre-existing)
    [runner-container-hooks] FATAL: Error: execPodStepWithRetry(mkdir temp dirs) failed after 6 attempts: Internal error occurred: unable to upgrade connection: container not found ("job")

CANCELLED JOBS - The following jobs were cancelled. Please retry:

This comment was automatically generated by Dr. CI and updates every 15 minutes.

@meta-cla meta-cla Bot added the CLA Signed This label is managed by the Facebook bot. Authors need to sign the CLA before a PR can be reviewed. label Sep 25, 2026
@github-actions

Copy link
Copy Markdown
Contributor

⚠️ PR Title Label Error

PR title must start with a label prefix in brackets (e.g., [BugFix]).

Current title: Move the Linux CI jobs to OSDC runners

Supported Prefixes (case-sensitive)

Your PR title must start with exactly one of these prefixes:

Prefix Label Applied Example
[Algorithm] new algo [Algorithm] Add new RL objective
[BE] BE [BE] Improve error messages
[Benchmark] or [Benchmarks] Benchmarks [Benchmark] Add collector benchmark
[BugFix] BugFix [BugFix] Fix memory leak in collector
[Example] or [Examples] Examples [Example] Add training script
[Feature] Feature [Feature] Add new optimizer
[Doc] or [Docs] Documentation [Doc] Update installation guide
[Refactor] Refactoring [Refactor] Clean up module imports
[CI] CI [CI] Fix workflow permissions
[Test] or [Tests] Tests [Tests] Add unit tests for buffer
[Trainer] or [Trainers] Trainers [Trainer] Add trainer config
[Environment] or [Environments] Environments [Environments] Add Gymnasium support
[Data] Data [Data] Fix replay buffer sampling
[LLM] llm/ [LLM] Add reward model integration
[Minor] small change [Minor] Fix typo in error message
[Performance] or [Perf] Performance [Performance] Optimize tensor ops
[BC-Breaking] bc breaking [BC-Breaking] Remove deprecated API
[Deprecation] Deprecation [Deprecation] Mark old function
[Algorithm] or [Algorithms] new algo [Algorithm] Add new objective
[Quality] Quality [Quality] Fix typos and add codespell
[Versioning] versioning [Versioning] Bump release version
[WIP] WIP [WIP] Draft implementation

Note: Common variations like singular/plural are supported (e.g., [Doc] or [Docs]).

@github-actions github-actions Bot added the CI Has to do with CI setup (e.g. wheels & builds, tests...) label Sep 25, 2026
@huydhn huydhn changed the title Move the Linux CI jobs to OSDC runners [BE] Move the Linux CI jobs to OSDC runners Oct 6, 2026
@github-actions github-actions Bot added the BE Better errors, logs, docs or test utils label Oct 6, 2026
@huydhn huydhn changed the title [BE] Move the Linux CI jobs to OSDC runners [Refactor] Move the Linux CI jobs to OSDC runners Oct 6, 2026
@huydhn huydhn added Refactoring Refactoring of an existing feature and removed BE Better errors, logs, docs or test utils labels Oct 6, 2026
huydhn added a commit that referenced this pull request Oct 6, 2026
The last linux_job.yml (v1) caller in this repo, and the remaining blocker
on pytorch/test-infra#8417. #4486 moved the rest of docs.yml to v3 and left
this one behind.

rsync becomes cp -a: v1's conda-builder image had rsync, almalinux-builder
(the v3 default) does not, and a plain recursive copy is all this needed.
Same swap as pytorch/tensordict#1807.

Authored with Claude Code.
huydhn added a commit that referenced this pull request Oct 6, 2026
Covers both of docs.yml's test-infra callers, so the whole file lands in
one change:

- build-docs: linux_job_v2 -> v3, linux.g5.4xlarge.nvidia.gpu ->
  mt-l-x86aavx2-11-41-a10g. Split out of #4486, which no longer touches
  this file.
- upload: linux_job.yml (v1) -> v3. The last v1 caller in the repo and the
  remaining blocker on pytorch/test-infra#8417.

rsync becomes cp -a: v1's conda-builder image had rsync, almalinux-builder
(the v3 default) does not, and a plain recursive copy is all this needed.
Same swap as pytorch/tensordict#1807.

Note upload only runs on a push to main or a tag, so CI on this PR does not
exercise it.

Authored with Claude Code.
Moves the Linux test and lint jobs across 11 workflows onto OSDC pods.

| before | after |
|---|---|
| `linux.g5.4xlarge.nvidia.gpu` | `mt-l-x86aavx2-11-41-a10g` |
| `linux.g6.4xlarge.experimental.nvidia.gpu` | `mt-l-x86aavx2-11-41-l4` |
| `linux.12xlarge` | `mt-l-x86iavx512-48-384` |
| `linux.4xlarge` | `mt-l-x86iavx512-16-128` |

Each pod sits on the same instance the label it replaces resolved to, so GPU
model and count are unchanged and only the Kubernetes overhead comes off.

`linux_job_v2` -> `v3` comes with it, per job rather than per file: OSDC
rejects a job that has no container, and v3 supplies
`pytorch/almalinux-builder` for the arch. Only jobs whose runner actually
moves are converted.

The lint jobs set no `runner:` at all, so they were quietly taking v2's
`linux.2xlarge` default. They now pin an OSDC pod explicitly.

Six setup scripts installed the EGL vendor config with

    cp $this_dir/10_nvidia.json /usr/share/glvnd/egl_vendor.d/10_nvidia.json

which dies on OSDC with "Read-only file system": the NVIDIA container
runtime bind-mounts that path. libglvnd searches
`<sysconfdir>/glvnd/egl_vendor.d` before `<datadir>/glvnd/egl_vendor.d`
(src/EGL/meson.build), so /etc/glvnd/egl_vendor.d is both writable and
higher priority. The scripts now install there, and skip when the runtime
already supplied the file.

docs.yml is handled in #4524 instead, which covers both of its callers --
build-docs here and the v1 gh-pages upload -- so the file lands in one
change rather than split across two PRs.

test-linux-libs.yml's `unittests-isaaclab` is deliberately left on EC2: it
calls `vmoens/test-infra/.../isaac_linux_job_v2.yml`, a fork's reusable
workflow that cannot be confirmed to set a container, and a container-less
job is rejected outright on OSDC.

Wheel builds are in a separate PR.

Authored with Claude Code.
@huydhn
huydhn force-pushed the osdc/ci-runners branch 4 times, most recently from e77f645 to e3707fa Compare October 7, 2026 02:31
Moves the Linux test and lint jobs onto OSDC pods.

| before | after |
|---|---|
| `linux.g5.4xlarge.nvidia.gpu` | `mt-l-x86aavx2-11-41-a10g` |
| `linux.g6.4xlarge.experimental.nvidia.gpu` | `mt-l-x86aavx2-11-41-l4` |
| `linux.12xlarge` | `mt-l-x86iavx512-48-384` |
| `linux.4xlarge` | `mt-l-x86iavx512-16-128` |

Each pod sits on the same instance the label it replaces resolved to, so GPU
model and count are unchanged and only the Kubernetes overhead comes off.

`linux_job_v2` -> `v3` comes with it, per job rather than per file: OSDC
rejects a job that has no container, and v3 supplies
`pytorch/almalinux-builder` for the arch.

The lint jobs set no `runner:` at all, so they were quietly taking v2's
`linux.2xlarge` default. They now pin an OSDC pod explicitly.

## Images

On OSDC the checkout runs inside the job container, and no stock nvidia/cuda
image ships git -- actions/checkout then falls back to a REST API tarball
with no .git, and the 88 unittest scripts that start with
`git rev-parse --show-toplevel` die. So 42 jobs move to the images
test-infra publishes for this, which are the same bases plus git:

  nvidia/cuda:12.4.0-devel-ubuntu22.04        -> osdc-cuda:cuda12.4.1-cudnn-devel-ubuntu22.04
  nvidia/cuda:13.0.2-cudnn-devel-ubuntu24.04  -> osdc-cuda:cuda13.0.3-cudnn-devel-ubuntu24.04
  nvidia/cuda:12.8.0-devel-ubuntu22.04        -> osdc-cuda:cuda12.8.1-cudnn-devel-ubuntu24.04
  nvidia/cuda:12.8.1-cudnn-devel-ubuntu22.04  -> osdc-cuda:cuda12.8.1-cudnn-devel-ubuntu24.04
  nvidia/cuda:11.8.0-cudnn8-devel-ubuntu22.04 -> osdc-cuda:cuda11.8.0-cudnn8-devel-ubuntu22.04

Three jobs have no equivalent yet and keep their image: unittests-robohive
on nvidia/cudagl, and unittests-vllm/sglang on pytorch/pytorch. None of
those carry git either, so they will need one adding to the test-infra
matrix.

## Also

Six setup scripts installed the EGL vendor config with

    cp $this_dir/10_nvidia.json /usr/share/glvnd/egl_vendor.d/10_nvidia.json

which dies on OSDC with "Read-only file system": the NVIDIA container
runtime bind-mounts that path. libglvnd searches
`<sysconfdir>/glvnd/egl_vendor.d` first, so the scripts now install to
/etc/glvnd/egl_vendor.d and skip when the runtime already supplied the file.

docs.yml is handled in #4524. test-linux-libs.yml's `unittests-isaaclab`
stays on EC2: it calls a fork's reusable workflow that cannot be confirmed
to set a container. Wheel builds are in a separate PR.

Authored with Claude Code.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

CI Has to do with CI setup (e.g. wheels & builds, tests...) CLA Signed This label is managed by the Facebook bot. Authors need to sign the CLA before a PR can be reviewed. Refactoring Refactoring of an existing feature

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant