DACE RL for Compute Efficient Reinforcement Learning in Small Model Reasoning
Abstract
Reinforcement learning with verifiable rewards (RLVR) has become a dominant recipe for improving mathematical and code reasoning in open-weight language models, but existing methods rely on fixed hyperparameters despite highly non-stationary training dynamics. We present DACE-RL (Diagnosis-Aware Compute-Efficient RL), a closed-loop framework that promotes training diagnostics to first-class signals. DACE-RL tracks four lightweight statistics—token entropy, sample diversity, advantage saturation, and answer consistency—and feeds their EMA-smoothed state to a controller that adaptively sets per-prompt rollout budget, sampling temperature, asymmetric clipping bounds, and entropy regularization strength. We formalize adaptive rollout allocation as a difficulty-conditioned multi-armed bandit and provide a regret bound that improves over uniform-budget GRPO under heterogeneous prompt difficulty; we also give a stability result for the dual-loop entropy-diversity controller. Across three small base models (Qwen2.5-Math-1.5B, Qwen2.5-7B, and Llama-3.1-8B-Instruct) and six math reasoning benchmarks, DACE-RL matches or improves Pass@1 while reducing rollout-GPU-hours by 38–52%, and improves Pass@32 by 4.7 absolute points on average. Pilot analysis further shows that diagnostic inflection signals appear 100–300 update steps before reward plateau, indicating that diagnosis-aware control is a key lever for compute-efficient RLVR.