Skip to content

Add CPU-side Kubernetes image-readiness preflight for VERL training #567

Description

Problem

In split CPU/GPU deployments, VERL can start and allocate GPUs before knowing whether the CPU-side Kubernetes nodes contain the images required by each dataset row. *This could casue hang for a long time for training **

SWE-Smith :openai images are built and tagged locally by pull_images.py; they are not pushed to a registry. The Job template uses imagePullPolicy:
IfNotPresent, so a missing local image can leave the Pod in image-pull backoff until the configured 1,800-second rollout deadline.

Failure mode when an image is missing

The failure is not detected before training starts:

  1. Ray initializes and reserves all 8 GPUs.
  2. The Controller creates rollout Jobs on the CPU Kubernetes node.
  3. Because the required :openai image is not present locally, Kubernetes attempts to pull it due to imagePullPolicy: IfNotPresent.
  4. These images are built locally and are not pushed to a registry, so the Pod remains Pending, typically with ErrImagePull or
    ImagePullBackOff.
  5. The rollout produces no trajectory, while the trainer continues waiting for the batch to complete.

**6. The failure may only surface at the configured 1,800-second rollout deadline—up to 30 minutes per affected batch. **

With train_batch_size=8 and rollout.n=8, one unavailable dataset row can create eight blocked rollout Jobs. This makes the problem appear as a
long training hang while allocated GPUs remain idle or underutilized.

In our readiness scan, 190/6,343 training rows and 126/474 validation rows required unavailable images. Without preflight filtering, all of these rows could enter this failure path.

Experiment configuration

model: Qwen/Qwen3.5-9B
topology: 8-GPU trainer + 1 CPU Kubernetes node
train_rows: 6343
validation_rows: 474

trainer.n_gpus_per_node: 8
trainer.total_training_steps: 800
trainer.val_before_train: false
trainer.test_freq: 100

data.train_batch_size: 8
actor_rollout_ref.rollout.n: 8
actor_rollout_ref.actor.ppo_mini_batch_size: 8
actor_rollout_ref.rollout.gpu_memory_utilization: 0.50

agentlightning.async_rollout.enabled: true
agentlightning.async_rollout.async_train_batch_size: 16
agentlightning.rollout_timeout_seconds: 1800

Rendered image pattern:

image: "{{ input.image_name }}:openai"
imagePullPolicy: IfNotPresent

## Proposal

Add an opt-in agentlightning.k8s.filter_unavailable_images setting, disabled by default.

The Controller publishes a leased CPU-node image inventory. Before ray.init(), the trainer renders the actual Job template and filters unavailable
train and validation rows. The Controller rechecks readiness immediately before creating guarded Jobs.

This changes no dataset files and transfers no images.

## Prototype validation

- Image-readiness tests: 51 passed
- Full test suite: 140 passed
- Disabled mode performs no readiness requests and preserves existing behavior

Would the maintainers accept this in core, or prefer it under contrib/?

Activity

  1. hzy46 commented on Aug 28, 2026

    @hzy46
    Contributor

    Hi baibizhe (SII) (@baibizhe) , thanks for reporting this issue.

    As far as I am concerned, the images and pods should not be managed by Agent Lightning. Instead, they should be managed by the Kubernetes. Thus, if the job fails because of image issues, the rollout in Agent Lightning is expected to be failed.

    The suitable way for handling this may be add some description in the document.

  2. baibizhe commented on Aug 29, 2026

    @baibizhe
    Author

    Thank you for your reply. Let me clarify the issue: when training is split between the CPU and GPU, the GPU doesn't know at startup which images on the CPU side are ready and which are not. It only finds this out when the actual training begins, which causes long hangs during the training process.

    Therefore, I have designed a solution: before the training officially starts, the GPU will check which images on the CPU side are good. If the images are not ready, all of that data will simply be filtered.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions