Repository navigation
[Refactor] Bound GPU test parallelism and reap strays - #4532
Merged
Merged
Conversation
huydhn
added this pull request to stack #4533
October 8, 2026 02:25
🔗 Helpful Links🧪 See artifacts and rendered test results at hud.pytorch.org/pr/pytorch/rl/4532
Note: Links to docs will display an error until the docs builds have been completed. ✅ You can merge normally! (6 Unrelated Failures)As of commit 133b21d with merge base 3231aea ( BROKEN TRUNK - The following jobs failed but were present on the merge base:👉 Rebase onto the `viable/strict` branch to avoid these failures
This comment was automatically generated by Dr. CI and updates every 15 minutes. |
This was referenced Oct 8, 2026
huydhn
removed this pull request from stack #4533
October 8, 2026 02:30
huydhn
force-pushed
the
osdc/test-parallelism
branch
from
October 8, 2026 02:31
b72e87b to
0674cac
Compare
huydhn
added this pull request to stack #4534
October 8, 2026 02:31
huydhn
force-pushed
the
osdc/test-parallelism
branch
from
October 8, 2026 07:15
0674cac to
c73c80b
Compare
huydhn
marked this pull request as ready for review
October 8, 2026 19:03
atalman
approved these changes
Oct 8, 2026
Two ways shard 3 failed on OSDC that have nothing to do with the tests. The job hung to its 120-minute timeout after the script had already exited 0. run_with_env_secrets.py drains the step's stdout until EOF, and EOF only arrives once every process holding the write end has closed it. A ps dump named the holder: Xvfb, started by the test setup and outliving pytest. Under linux_job_v2 the surrounding `docker run` reaped it for us. So dump the surviving processes and reap them before exiting. The dump also shows Xvfb outliving pytest on the EC2 runners, so this is not OSDC-specific. -n auto gives 16 workers on the GPU pods, and 16 concurrent CUDA contexts on one A10G leave no room: a test that spawns its own CUDA subprocess dies in cuDevicePrimaryCtxRetain with CUDA_ERROR_OUT_OF_MEMORY before allocating anything, and ordinary Triton tests start failing to allocate too. Which tests lose the race varies run to run, so cap GPU concurrency at 4 rather than chase them. Shard 3 goes from ~650s to ~1130s, well inside the timeout. The CPU path keeps its measured 24. Authored with Claude Code.
huydhn
force-pushed
the
osdc/test-parallelism
branch
from
October 9, 2026 01:21
c73c80b to
133b21d
Compare
Contributor
Author
|
@pytorchbot drci |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Second of three. Two ways the test run misbehaves in a job container, both independent of where the job runs.
run_with_env_secrets.pydrains the step's stdout until EOF, and EOF waits on every process holding the write end. Apsdump named the holder: Xvfb, started by the test setup and outliving pytest. Underlinux_job_v2the surroundingdocker runreaped it. Now the script dumps survivors and reaps them.-n autogives 16 xdist workers on the GPU pods. 16 concurrent CUDA contexts on one A10G leave no room: a test spawning its own CUDA subprocess dies incuDevicePrimaryCtxRetainwithCUDA_ERROR_OUT_OF_MEMORY, and ordinary Triton tests fail to allocate. Which tests lose varies per run, so cap GPU concurrency at 4. Shard 3 goes ~650s → ~1130s, well inside the timeout. CPU keeps its measured 24.test_dreamer_v3_checkpoint_resume_processesis also deselected from the xdist runs and added to the serial shard — the existing quarantine is path-based, and moving all oftest_dreamer_v3.pywould cost ~1300s of the shard's ~2950s.Stacked on #4531, below #4530.
Authored with Claude Code.