$R_Q$-Evolve: Evolving Task Generators at the Learning Frontier of Self-Evolving LLMs
Yonghoon Kwon ⋅ Hoyoon Byun ⋅ Kyungwoo Song
Abstract
The learning frontier moves as the Solver improves, so a self-evolving language model must continually adapt its training problems. Learned task proposers encode this adaptation implicitly in their parameters. Previously too-difficult discoveries can become informative, while informative ones can become uniformly solvable. An adaptive curriculum therefore needs not only to generate new problems, but also to preserve and reassess previous discoveries. We introduce $R_Q$-Evolve, which makes retained discoveries explicit in a persistent archive of generator programs. Preserving generation rules allows task families to be re-evaluated through fresh instances and structurally mutated. Solver-conditioned scoring combines learnability and token entropy to guide selection, MAP-Elites preservation retains a representative per mathematical domain and output form so no region crowds out the rest, and two-stage mutation implements expansion of this persistent curriculum state without a separately optimized task proposer. On Qwen3-4B-Base and Qwen3-8B-Base, $R_Q$-Evolve raises mathematical benchmark averages by $5.24$ and $6.49$ points over the base models and exceeds the compared self-evolving baselines. Generated 8B instances stay broadly distributed late in training, and the complete run is $3.39\times$ faster end-to-end than the Challenger-based R-Zero at a matched Solver update budget.
Chat is not available.
Successful Page Load