EfficientRollout: System-Aware Self-Speculative Decoding for RL Rollouts
Abstract
Reinforcement learning (RL) has become a representative post-training paradigm for large language models (LLMs), enabling strong reasoning and agentic capabilities. However, its rollout generation remains a dominant training bottleneck because it relies on sequential autoregressive (AR) decoding, where a small number of long-tailed responses often determine completion time. Speculative decoding (SD) can reduce inference latency while preserving model quality by rapidly drafting tokens and accepting them through parallel verification. Applying SD to RL rollouts, however, introduces challenges absent from standard LLM inference: (i) algorithmically, the continuously evolving target model makes static drafters stale; (ii) system-wise, rollout decoding moves across regimes, from large active batches where SD can be compute-bound and ineffective to shrinking-batch tails where SD becomes beneficial. Existing RL-SD approaches address aspects of this problem, but either yield low effective accepted lengths that limit speedup or rely on auxiliary drafters requiring pretraining and online adaptation, increasing system complexity. We present EfficientRollout, an SD framework designed to accelerate RL rollouts while addressing these challenges. EfficientRollout induces a quantized drafter directly from the target model, keeping it coupled to the evolving policy without separate drafter training before or during RL. It then uses a dynamic SD toggle policy that enables SD only in beneficial regimes identified by system-aware roofline modeling. It further adapts drafting budgets using acceptance-behavior signals observed during training, better realizing the potential accepted length. Under realistic RL workload, EfficientRollout reduces rollout and end-to-end latency by up to 19.6% and 12.7%, respectively, over a standard accelerated AR rollout baseline, while preserving final model quality.