Adaptive Scheduling Pipeline For Multi-Instance Asynchronous Reinforcement Learning
Salah Chikhi ⋅ Luis H Ruiz ⋅ Entong Li ⋅ Li Zeng
Abstract
We present the Adaptive Scheduling Pipeline (ASP), a bounded-lag asynchronous reinforcement learning (RL) architecture for post-training large language models (LLMs). Rather than only disaggregating generation and training, the ASP combines two design ideas. First, a queue-aware scheduling policy that jointly reduces three sources of rollout underutilization —padding bubbles from intra-minibatch length heterogeneity, skewness bubbles from long-tail outputs, and sorting bubbles from queue instability induced by length-sorted dispatching—through a single offline preprocessing pass. This enables efficient multi-instance rollout that overcomes the early scaling plateau of single-instance inference while maintaining stable queue dynamics. Second, a bounded-lag asynchronous architecture decouples rollout from training while keeping policy lag observable and small enough to preserve stable learning in our experiments. We also provide an explicit weight-flow path that converts trainer snapshots into inference-ready weights without re-coupling generation and training. Experiments show that ASP improves end-to-end throughput by up to $3.22\times$, reduces iteration time by up to 76\%, scales more robustly with additional generation resources than the MindSpeed-RL and verl baselines, and shows no visible convergence degradation. ASP has also been used in an internal telecom-domain RL post-training workflow, reducing the reported training period from more than three months to about two months and producing a domain-enhanced model for knowledge answering and tool calling.
Chat is not available.
Successful Page Load