Train at the Moving Edge: Rollout-Efficient RL for Large Reasoning Models
Abstract
Reinforcement learning (RL) has become essential for post-training large language models (LLMs) in reasoning tasks. While scaling rollouts can stabilize training and enhance performance, it introduces substantial computational overhead. In algorithms like GRPO, multiple rollouts per prompt incur prohibitive costs, as a large portion of prompts provide negligible gradients and are thus of low utility. This raises a key question: \textit{how to identify high-utility prompts before an expensive rollout?} Our experimental analysis reveals that sample utility is non-uniform and dynamic: the strongest learning signals concentrate at the ``learning edge'', the intersection of intermediate difficulty and high response entropy, which shifts throughout training. Motivated by this observation, we propose HIVE, a history-informed and online-verified prompt selection framework for data-efficient RL training. HIVE first uses historical reward statistics and response entropy as a cheap prior to filter candidate prompts, and then employs prompt entropy as a real-time proxy to prune instances with stale utility. Across multiple reasoning benchmarks and base models, HIVE maintains reasoning accuracy with substantial rollout cost reduction, achieving up to 2.3× total training speedup.