Skip to content

[Refactor] Bound GPU test parallelism and reap strays - #4532

Merged
huydhn merged 1 commit into
mainfrom
osdc/test-parallelism
Oct 9, 2026
Merged

huydhn merged 1 commit into
mainfrom
osdc/test-parallelism

Conversation

@huydhn

@huydhn huydhn commented Oct 8, 2026 •

Copy link
Copy Markdown
Contributor

Second of three. Two ways the test run misbehaves in a job container, both independent of where the job runs.

  • Jobs hung to the 120-minute timeout after the script had already exited 0. run_with_env_secrets.py drains the step's stdout until EOF, and EOF waits on every process holding the write end. A ps dump named the holder: Xvfb, started by the test setup and outliving pytest. Under linux_job_v2 the surrounding docker run reaped it. Now the script dumps survivors and reaps them.
  • -n auto gives 16 xdist workers on the GPU pods. 16 concurrent CUDA contexts on one A10G leave no room: a test spawning its own CUDA subprocess dies in cuDevicePrimaryCtxRetain with CUDA_ERROR_OUT_OF_MEMORY, and ordinary Triton tests fail to allocate. Which tests lose varies per run, so cap GPU concurrency at 4. Shard 3 goes ~650s → ~1130s, well inside the timeout. CPU keeps its measured 24.

test_dreamer_v3_checkpoint_resume_processes is also deselected from the xdist runs and added to the serial shard — the existing quarantine is path-based, and moving all of test_dreamer_v3.py would cost ~1300s of the shard's ~2950s.

Stacked on #4531, below #4530.

Authored with Claude Code.

@huydhn
huydhn added this pull request to stack #4533 October 8, 2026 02:25
@pytorch-bot

pytorch-bot Bot commented Oct 8, 2026 •

Copy link
Copy Markdown

🔗 Helpful Links

🧪 See artifacts and rendered test results at hud.pytorch.org/pr/pytorch/rl/4532

Note: Links to docs will display an error until the docs builds have been completed.

✅ You can merge normally! (6 Unrelated Failures)

As of commit 133b21d with merge base 3231aea (image):

BROKEN TRUNK - The following jobs failed but were present on the merge base:

👉 Rebase onto the `viable/strict` branch to avoid these failures

This comment was automatically generated by Dr. CI and updates every 15 minutes.

@meta-cla meta-cla Bot added the CLA Signed This label is managed by the Facebook bot. Authors need to sign the CLA before a PR can be reviewed. label Oct 8, 2026
@github-actions github-actions Bot added CI Has to do with CI setup (e.g. wheels & builds, tests...) Refactoring Refactoring of an existing feature labels Oct 8, 2026
@huydhn
huydhn removed this pull request from stack #4533 October 8, 2026 02:30
@huydhn
huydhn force-pushed the osdc/test-parallelism branch from b72e87b to 0674cac Compare October 8, 2026 02:31
@huydhn
huydhn added this pull request to stack #4534 October 8, 2026 02:31
@huydhn
huydhn force-pushed the osdc/test-parallelism branch from 0674cac to c73c80b Compare October 8, 2026 07:15
@huydhn
huydhn requested a review from atalman October 8, 2026 18:45
@huydhn
huydhn marked this pull request as ready for review October 8, 2026 19:03
Base automatically changed from osdc/setup-scripts to main October 9, 2026 01:21
Two ways shard 3 failed on OSDC that have nothing to do with the tests.

The job hung to its 120-minute timeout after the script had already exited
0. run_with_env_secrets.py drains the step's stdout until EOF, and EOF only
arrives once every process holding the write end has closed it. A ps dump
named the holder: Xvfb, started by the test setup and outliving pytest.
Under linux_job_v2 the surrounding `docker run` reaped it for us. So dump
the surviving processes and reap them before exiting. The dump also shows
Xvfb outliving pytest on the EC2 runners, so this is not OSDC-specific.

-n auto gives 16 workers on the GPU pods, and 16 concurrent CUDA contexts on
one A10G leave no room: a test that spawns its own CUDA subprocess dies in
cuDevicePrimaryCtxRetain with CUDA_ERROR_OUT_OF_MEMORY before allocating
anything, and ordinary Triton tests start failing to allocate too. Which
tests lose the race varies run to run, so cap GPU concurrency at 4 rather
than chase them. Shard 3 goes from ~650s to ~1130s, well inside the
timeout. The CPU path keeps its measured 24.

Authored with Claude Code.
@huydhn
huydhn force-pushed the osdc/test-parallelism branch from c73c80b to 133b21d Compare October 9, 2026 01:21
@huydhn

huydhn commented Oct 9, 2026

Copy link
Copy Markdown
Contributor Author

@pytorchbot drci

@huydhn
huydhn merged commit 3a13e5f into main Oct 9, 2026
122 of 128 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

CI Has to do with CI setup (e.g. wheels & builds, tests...) CLA Signed This label is managed by the Facebook bot. Authors need to sign the CLA before a PR can be reviewed. Refactoring Refactoring of an existing feature

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants