LADDERS: Length-Aware Data Distribution and Existing-Response Speculation for Fast RL Rollout Generation
Abstract
Reinforcement learning (RL) has become central to post-training large language models, but its rollout generation stage is often the dominant system bottleneck. The inefficiency comes from two sources. First, autoregressive decoding makes latency grow with response length. Second, response lengths vary widely within a batch: short sequences finish early, but the batch remains blocked by the longest sequence, leaving GPU capacity underutilized. This effect is amplified in multi-sample rollouts, where long responses can also concentrate on a few decoding workers and delay synchronization. We propose LADDERS (Length-Aware Data Distribution and Existing-Response Speculation), a lightweight framework for accelerating RL rollout generation. LADDERS first uses hidden-state-based length prediction to group prompts with similar expected response lengths, and then applies an S-shaped allocation rule to balance multi-sample requests across workers. It further reuses responses already generated by the policy as draft continuations through prompt-specific suffix trees, enabling speculative decoding without an auxiliary draft model. Experiments on Qwen3-1.7B and Qwen3-8B show that LADDERS reduces rollout generation time by up to 56% while preserving final task performance, and can be integrated into existing RL systems with minimal engineering effort.