From Reasoning Chains to Verifiable Subproblems: Curriculum Reinforcement Learning Enables Credit Assignment for LLM Reasoning
Xitai Jiang ⋅ Wenze Lin ⋅ Zihan Tang ⋅ Yang Yue ⋅ Shenzhi Wang ⋅ Gao Huang
Abstract
Reinforcement learning from verifiable rewards (RLVR) has shown strong promise for LLM reasoning, but typical outcome-based RLVR methods remain inefficient on hard problems. Correct final-answer rollouts are rare, and standard sample-level credit assignment fails to leverage partial reasoning progress embedded in unsuccessful attempts. To address this problem, we introduce **SCRL** (**S**ubproblem **C**urriculum **R**einforcement **L**earning), a curriculum reinforcement learning framework built on verifiable subproblems derived from reasoning chains. Given a reference solution, SCRL derives a series of verifiable subproblems and constructs a subproblem curriculum, with the final subproblem fixed as the original problem. This converts partial progress on hard problems into verifiable learning signals. Algorithmically, we propose *subproblem-level normalization*, a training technique based on RLVR that normalizes rewards independently at each subproblem position within the rollout group. By assigning the resulting advantages to the corresponding answer spans, we enable finer-grained credit assignment without external rubrics or reward models. Our theoretical analysis shows that this subproblem curriculum makes hard problems more learnable by lifting them out of gradient dead zones, with larger relative gains as the original problem becomes harder. Across seven mathematical reasoning benchmarks, SCRL outperforms strong curriculum-learning baselines, yielding +4.1 and +1.9 average-point gains compared to GRPO on Qwen3-4B-Base and Qwen3-14B-Base respectively. On three hard benchmarks (AIME24, AIME25, and IMO-Bench), SCRL further yields point gains of +3.7 in pass@$1$ and +4.6 in pass@$64$ on Qwen3-4B-Base, suggesting improved exploration on hard reasoning problems. Further ablations show that the proposed credit assignment is effective and that the gains do not require highly curated subproblems or strong external generators, highlighting SCRL as a practical curriculum framework that enables fine-grained credit assignment for LLM reasoning.
Chat is not available.
Successful Page Load