Skip to content

[Refactor] Bound GPU test parallelism and reap strays - #4529

Closed
huydhn wants to merge 1 commit into
gh/huydhn/3/basefrom
gh/huydhn/3/head
Closed

huydhn wants to merge 1 commit into
gh/huydhn/3/basefrom
gh/huydhn/3/head

Conversation

@huydhn

@huydhn huydhn commented Oct 8, 2026 •

Copy link
Copy Markdown
Contributor

Stack from ghstack (oldest at bottom):

Two ways shard 3 failed on OSDC that have nothing to do with the tests.

The job hung to its 120-minute timeout after the script had already exited
0. run_with_env_secrets.py drains the step's stdout until EOF, and EOF only
arrives once every process holding the write end has closed it. A ps dump
named the holder: Xvfb, started by the test setup and outliving pytest.
Under linux_job_v2 the surrounding docker run reaped it for us. So dump
the surviving processes and reap them before exiting.

-n auto gives 16 workers on the GPU pods, and 16 concurrent CUDA contexts on
one A10G leave no room: a test that spawns its own CUDA subprocess dies in
cuDevicePrimaryCtxRetain with CUDA_ERROR_OUT_OF_MEMORY before allocating
anything, and ordinary Triton tests start failing to allocate too. Which
tests lose the race varies run to run, so cap GPU concurrency at 4 rather
than chase them. Shard 3 goes from ~650s to ~1130s, well inside the timeout.
The CPU path keeps its measured 24.

test_dreamer_v3_checkpoint_resume_processes is also deselected from the
xdist runs and added to the serial shard. The existing quarantine is
path-based and moving all of test_dreamer_v3.py would cost ~1300s of the
shard's ~2950s, so quarantine just the one test.

Authored with Claude Code.

[ghstack-poisoned]
@pytorch-bot

pytorch-bot Bot commented Oct 8, 2026

Copy link
Copy Markdown

🔗 Helpful Links

🧪 See artifacts and rendered test results at hud.pytorch.org/pr/pytorch/rl/4529

Note: Links to docs will display an error until the docs builds have been completed.

This comment was automatically generated by Dr. CI and updates every 15 minutes.

@meta-cla meta-cla Bot added the CLA Signed This label is managed by the Facebook bot. Authors need to sign the CLA before a PR can be reviewed. label Oct 8, 2026
@github-actions github-actions Bot added Refactoring Refactoring of an existing feature CI Has to do with CI setup (e.g. wheels & builds, tests...) and removed Refactoring Refactoring of an existing feature labels Oct 8, 2026
@huydhn

huydhn commented Oct 8, 2026

Copy link
Copy Markdown
Contributor Author

(Claude, on huydhn's behalf.) Superseded by the gh stack version of the same split: #4530 / #4531 / #4532. Same three commits, same content — this one was built with ghstack by mistake.

@huydhn huydhn closed this Oct 8, 2026
huydhn added a commit that referenced this pull request Oct 8, 2026
Two ways shard 3 failed on OSDC that have nothing to do with the tests.

The job hung to its 120-minute timeout after the script had already exited
0. run_with_env_secrets.py drains the step's stdout until EOF, and EOF only
arrives once every process holding the write end has closed it. A ps dump
named the holder: Xvfb, started by the test setup and outliving pytest.
Under linux_job_v2 the surrounding `docker run` reaped it for us. So dump
the surviving processes and reap them before exiting.

-n auto gives 16 workers on the GPU pods, and 16 concurrent CUDA contexts on
one A10G leave no room: a test that spawns its own CUDA subprocess dies in
cuDevicePrimaryCtxRetain with CUDA_ERROR_OUT_OF_MEMORY before allocating
anything, and ordinary Triton tests start failing to allocate too. Which
tests lose the race varies run to run, so cap GPU concurrency at 4 rather
than chase them. Shard 3 goes from ~650s to ~1130s, well inside the timeout.
The CPU path keeps its measured 24.

test_dreamer_v3_checkpoint_resume_processes is also deselected from the
xdist runs and added to the serial shard. The existing quarantine is
path-based and moving all of test_dreamer_v3.py would cost ~1300s of the
shard's ~2950s, so quarantine just the one test.

Authored with Claude Code.


ghstack-source-id: 5bf4013
Pull-Request: #4529
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

CI Has to do with CI setup (e.g. wheels & builds, tests...) CLA Signed This label is managed by the Facebook bot. Authors need to sign the CLA before a PR can be reviewed. Refactoring Refactoring of an existing feature

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant