Skip to content

perf(verl): batch rollout status polling - #612

Open
Bozhen Peng (kiteretsu903) wants to merge 1 commit into
microsoft:mainfrom
kiteretsu903:perf-batch-rollout-status
Open

Bozhen Peng (kiteretsu903) wants to merge 1 commit into
microsoft:mainfrom
kiteretsu903:perf-batch-rollout-status

Conversation

@kiteretsu903

Copy link
Copy Markdown
Contributor

perf(verl): batch rollout status polling

I noticed that both rollout managers fetch the full rollout, including its input and config, every time they check whether an agent has finished. The Search-R1 example creates 2,048 rollouts per training batch, so that can mean 2,048 separate HTTP requests in a single sweep, even when everyone is still running.

This adds a read-only POST /api/rollouts/status endpoint and uses it to check up to 256 rollouts at a time. Full details are fetched when a rollout finishes, so completion hooks still receive the input and metadata they need. The status response keeps all lifecycle fields, including timestamps. I also kept transient retries while making permanent errors such as a missing rollout fail immediately.

In a local benchmark with 2,048 rollouts and synthetic 4 KB inputs, one status sweep went from 2,048 requests to 8, and response data fell from 9.20 MB to 0.40 MB. Median sweep time over three trials was 1.82 s before and 22 ms after.

The Gateway and trainer need to be upgraded together for the new endpoint. I documented the request format, batch limit, authentication, and error behavior.

Validation

I used Python 3.12 with the locked CPU dependencies and added focused coverage for batch boundaries, retries, lifecycle fields, hook rewards, timestamps, and asynchronous group carry-over.

  • python -m pytest -q --durations=10 tests — 145 passed.
  • python -m pytest -q tests/server tests/controller tests/test_package.py tests/examples/test_swe_smith_images.py — 64 passed in a separate environment with only the release workflow's core/dev dependencies.
  • ruff check ., ruff format --check ., python scripts/check_headers.py, and pre-commit run --all-files --show-diff-on-failure — passed.
  • pyright --venvpath /home/admin1/os-contributions/agent-lightning-research — no errors or warnings.
  • uv build --no-sources --out-dir /mnt/d/Documents/os-contributions/verification/agent-lightning-next/dist — passed; checked the updated modules in both archives.
  • mkdocs build --strict --site-dir /mnt/d/Documents/os-contributions/verification/agent-lightning-next/site — passed.

Copilot AI balanced review requested due to automatic review settings October 1, 2026 07:42

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot was unable to review this pull request because the user who requested the review has reached their quota limit.

@hzy46

Copy link
Copy Markdown
Contributor

It is useful but using "POST /api/rollouts/status" is strange for a read-only endpoint.

Maybe "GET /api/rollouts/statuses"?

@kiteretsu903

Bozhen Peng (kiteretsu903) commented Oct 9, 2026 •

Copy link
Copy Markdown
Contributor Author

It is useful but using "POST /api/rollouts/status" is strange for a read-only endpoint.

Maybe "GET /api/rollouts/statuses"?

Thanks for the feedback! The design is to check 256 rollouts at once, and their IDs can go beyond 8 KB proxy limit on GET. But POST puts the IDs in the request body to avoid this problem.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants