Reasoning Warm-up: Scaling Label-free RL via Verifiable Surrogate Rewards
Abstract
Improving the reasoning ability of large language models (LLMs) without ground-truth supervision remains a central challenge in reinforcement learning (RL). Recent label-free RL methods typically rely on internal pseudo-rewards such as self-consistency or majority voting. However, for complex multi-step reasoning, these signals can be systematically misleading: models may repeatedly produce mutually consistent but incorrect trajectories, causing optimization to favor frequency over correctness. In this work, we propose reasoning warm-up reinforcement learning, a label-free RL framework motivated by an empirical phenomenon we call Reasoning Warm-up. We observe that when a model successfully completes a deterministic auxiliary task with an exact verifier, its success probability on a subsequent complex reasoning task increases substantially. This coupling suggests that auxiliary-task verification can serve as a useful surrogate signal for selecting and optimizing reasoning trajectories even when the main task itself is not directly verifiable. Based on this observation, we incorporate verifiable auxiliary tasks into both generation and optimization. During generation, the auxiliary task provides verifiable warm-up prefix; during training, its verification outcome is used to score candidate trajectories for policy optimization. Experiments on multiple benchmarks demonstrate the effectiveness of our method.