Beyond difficulties: Insights for provably efficient design of autocurriculum in RLVR
Ruofan Wu ⋅ Xin Ge ⋅ Gao Xuemin ⋅ Guanhua Fang ⋅ Yangbin Shi ⋅ Dongbai Guo
Abstract
The design of an efficient curriculum has become increasingly important for reinforcement learning with verifiable rewards (RLVR). The majority of existing curriculum strategies select training problems based on **difficulty**---typically measured by success rate---yet difficulty is an indirect proxy that conflates problem hardness with actual training utility. In this paper, we propose to drive curriculum design not by how hard a problem is, but by how much the model improves from training on it. We formalize this principle through the **improvement function** (IF), which measures the gain in a prompt's expected reward under an infinitesimal GRPO policy gradient step. Leveraging the IF framework, we formally characterize the recently discovered **edge of competence** (EoC) phenomenon: a problem is at the model's EoC when its improvement is within a constant factor of the maximum over the training set. Analyzing GRPO on $L$-step compositional reasoning tasks, we prove that an EoC-induced curriculum achieves target mastery in $\widetilde{\Theta}(\log L_{\max} / (\eta \log d))$ steps---an exponential improvement over the $\widetilde{\Theta}(L_{\max} / (\eta \log d))$ steps required by a uniform mixture of difficulties. Motivated by this theory, we propose ReCUR, a practical algorithm that maintains a stratified retry buffer of previously failed problems, resampling them until they yield positive reward or exhaust a retry budget. Experiments on multimodal reasoning benchmarks show that ReCUR consistently improves the performance of GRPO and DAPO, providing empirical evidence consistent with the improvement-driven curriculum perspective.
Chat is not available.
Successful Page Load