Skip to content

[Refactor] Move the Linux CI jobs to OSDC runners - #4530

Merged
huydhn merged 6 commits into
mainfrom
osdc/runners
Oct 9, 2026
Merged

huydhn merged 6 commits into
mainfrom
osdc/runners

Conversation

@huydhn

@huydhn huydhn commented Oct 8, 2026 •

Copy link
Copy Markdown
Contributor

Last of three — the actual fleet move. Everything it depends on is already on main by this point.

before after
linux.g5.4xlarge.nvidia.gpu mt-l-x86aavx2-11-41-a10g
linux.g6.4xlarge.experimental.nvidia.gpu mt-l-x86aavx2-11-41-l4
linux.12xlarge mt-l-x86iavx512-48-384
linux.4xlarge mt-l-x86iavx512-16-128

Each pod sits on the same instance the label it replaces resolved to, so GPU model and count are unchanged. linux_job_v2 → v3 comes with it, per job: OSDC rejects a job with no container. The lint jobs set no runner: at all and were silently taking v2's linux.2xlarge default — now pinned explicitly.

Images. On OSDC the checkout runs inside the job container and no stock nvidia/cuda ships git, so 42 jobs move to the ghcr.io/pytorch/test-infra/osdc-cuda images — the same bases plus git.

Two exceptions. unittests-robohive keeps nvidia/cudagl:11.4.0-base, the only image with the EGL stack it renders with, so it runs standalone with an explicit git install. unittests-isaaclab stays on EC2 — it calls a fork's reusable workflow that cannot be confirmed to set a container.

docs.yml is #4524; wheel builds are #4487. Supersedes #4486.

Authored with Claude Code.

@pytorch-bot

pytorch-bot Bot commented Oct 8, 2026 •

Copy link
Copy Markdown

🔗 Helpful Links

🧪 See artifacts and rendered test results at hud.pytorch.org/pr/pytorch/rl/4530

Note: Links to docs will display an error until the docs builds have been completed.

❌ 4 New Failures, 2 Unrelated Failures

As of commit c1cfb1a with merge base 3a13e5f (image):

NEW FAILURES - The following jobs have failed:

FLAKY - The following job failed but was likely due to flakiness present on trunk:

BROKEN TRUNK - The following job failed but was present on the merge base:

👉 Rebase onto the `viable/strict` branch to avoid these failures

This comment was automatically generated by Dr. CI and updates every 15 minutes.

@meta-cla meta-cla Bot added the CLA Signed This label is managed by the Facebook bot. Authors need to sign the CLA before a PR can be reviewed. label Oct 8, 2026
@huydhn
huydhn added this pull request to stack #4533 October 8, 2026 02:25
@github-actions

github-actions Bot commented Oct 8, 2026

Copy link
Copy Markdown
Contributor

⚠️ PR Title Label Error

PR title must start with a label prefix in brackets (e.g., [BugFix]).

Current title: osdc/runners

Supported Prefixes (case-sensitive)

Your PR title must start with exactly one of these prefixes:

Prefix Label Applied Example
[Algorithm] new algo [Algorithm] Add new RL objective
[BE] BE [BE] Improve error messages
[Benchmark] or [Benchmarks] Benchmarks [Benchmark] Add collector benchmark
[BugFix] BugFix [BugFix] Fix memory leak in collector
[Example] or [Examples] Examples [Example] Add training script
[Feature] Feature [Feature] Add new optimizer
[Doc] or [Docs] Documentation [Doc] Update installation guide
[Refactor] Refactoring [Refactor] Clean up module imports
[CI] CI [CI] Fix workflow permissions
[Test] or [Tests] Tests [Tests] Add unit tests for buffer
[Trainer] or [Trainers] Trainers [Trainer] Add trainer config
[Environment] or [Environments] Environments [Environments] Add Gymnasium support
[Data] Data [Data] Fix replay buffer sampling
[LLM] llm/ [LLM] Add reward model integration
[Minor] small change [Minor] Fix typo in error message
[Performance] or [Perf] Performance [Performance] Optimize tensor ops
[BC-Breaking] bc breaking [BC-Breaking] Remove deprecated API
[Deprecation] Deprecation [Deprecation] Mark old function
[Algorithm] or [Algorithms] new algo [Algorithm] Add new objective
[Quality] Quality [Quality] Fix typos and add codespell
[Versioning] versioning [Versioning] Bump release version
[WIP] WIP [WIP] Draft implementation

Note: Common variations like singular/plural are supported (e.g., [Doc] or [Docs]).

@github-actions github-actions Bot added the CI Has to do with CI setup (e.g. wheels & builds, tests...) label Oct 8, 2026
@huydhn huydhn changed the title osdc/runners [Refactor] Move the Linux CI jobs to OSDC runners Oct 8, 2026
@github-actions github-actions Bot added the Refactoring Refactoring of an existing feature label Oct 8, 2026
@huydhn
huydhn removed this pull request from stack #4533 October 8, 2026 02:30
@huydhn
huydhn changed the base branch from main to osdc/test-parallelism October 8, 2026 02:31
@huydhn
huydhn added this pull request to stack #4534 October 8, 2026 02:31
@huydhn

huydhn commented Oct 8, 2026

Copy link
Copy Markdown
Contributor Author

@pytorchbot drci

@huydhn
huydhn requested a review from atalman October 8, 2026 19:07
@huydhn
huydhn marked this pull request as ready for review October 8, 2026 19:07
Base automatically changed from osdc/test-parallelism to main October 9, 2026 02:16
huydhn added 6 commits October 8, 2026 19:16
Moves the Linux test and lint jobs across 11 workflows onto OSDC pods.

| before | after |
|---|---|
| `linux.g5.4xlarge.nvidia.gpu` | `mt-l-x86aavx2-11-41-a10g` |
| `linux.g6.4xlarge.experimental.nvidia.gpu` | `mt-l-x86aavx2-11-41-l4` |
| `linux.12xlarge` | `mt-l-x86iavx512-48-384` |
| `linux.4xlarge` | `mt-l-x86iavx512-16-128` |

Each pod sits on the same instance the label it replaces resolved to, so GPU
model and count are unchanged and only the Kubernetes overhead comes off.
`linux_job_v2` -> `v3` comes with it, per job: OSDC rejects a job with no
container.

The lint jobs set no `runner:` at all, so they were quietly taking v2's
`linux.2xlarge` default. They now pin an OSDC pod explicitly.

On OSDC the checkout runs inside the job container and no stock nvidia/cuda
image ships git, so 42 jobs move to the images test-infra publishes for
this -- the same bases plus git:

  nvidia/cuda:12.4.0-devel-ubuntu22.04        -> osdc-cuda:cuda12.4.1-cudnn-devel-ubuntu22.04
  nvidia/cuda:13.0.2-cudnn-devel-ubuntu24.04  -> osdc-cuda:cuda13.0.3-cudnn-devel-ubuntu24.04
  nvidia/cuda:12.8.x-*-ubuntu22.04            -> osdc-cuda:cuda12.8.1-cudnn-devel-ubuntu24.04
  nvidia/cuda:11.8.0-cudnn8-devel-ubuntu22.04 -> osdc-cuda:cuda11.8.0-cudnn8-devel-ubuntu22.04
  pytorch/pytorch:2.8.0-cuda12.9-cudnn9-devel -> osdc-cuda:cuda12.9.2-cudnn-devel-ubuntu24.04

unittests-robohive keeps nvidia/cudagl:11.4.0-base -- it is the only image
providing the EGL stack that suite renders with -- so it runs standalone
with an explicit git install instead of through v3.

test-linux-libs.yml's `unittests-isaaclab` stays on EC2: it calls a fork's
reusable workflow that cannot be confirmed to set a container.

docs.yml is handled in #4524. Wheel builds are in #4487.

Authored with Claude Code.


ghstack-source-id: b70316e
Pull-Request: #4527
@huydhn
huydhn merged commit 47bd395 into main Oct 9, 2026
126 of 132 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

CI Has to do with CI setup (e.g. wheels & builds, tests...) CLA Signed This label is managed by the Facebook bot. Authors need to sign the CLA before a PR can be reviewed. Refactoring Refactoring of an existing feature

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants