Fast-RL: Accelerating Reinforcement Learning for LongCoT Reasoning Models
Sitong Wu ⋅ Haoru Tan ⋅ bin xia ⋅ Bei Yu ⋅ Xiaojuan Qi ⋅ Jiaya Jia
Abstract
Large Reasoning Models (LRMs) have made remarkable progress, driven by a paradigm shift from fast thinking to slow thinking. While slow-thinking unlocks superior reasoning ability, its long-form nature imposes a prohibitive computational burden on Reinforcement Learning (RL) for LRMs. In this paper, we propose **Fast-RL**, a novel framework that accelerates RL training for slow-thinking LRMs. Our acceleration principle stems from a foundational insight: fast-thinking sampling serves as an efficient and powerful proxy for slow-thinking sampling in enhancing reasoning, since the core reasoning skills are shared across thinking modes and can be improved through either trajectory type. Fast-RL consists of two stages. **Stage-1** establishes reliable prompt-based thinking-mode control, so that a slow-thinking-oriented model triggers fast-thinking under a specific prompt while preserving its native slow-thinking under the standard prompt. **Stage-2** then enhances reasoning ability through RL with exclusively fast-thinking sampling for acceleration. A lightweight LoRA adapter is introduced as a dedicated mode controller: it is trained in Stage-1 and **frozen** in Stage-2, decoupling mode-control parameters from reasoning parameters and protecting the model's native slow-thinking from being disturbed by fast-thinking RL. After Stage-2, the LoRA acts as a disposable training scaffold that can be either merged into the main model or simply discarded. Experiments on diverse multimodal and text-only LRMs show that Fast-RL accelerates training by about $10\times$ while delivering superior performance gains over vanilla slow-thinking RL. On Qwen3-1.7B, Fast-RL improves AIME24 by $+10.0$ and $+13.7$ in slow- and fast-thinking evaluation, respectively. Vanilla slow-thinking RL, in contrast, drops slow-thinking accuracy by $-8.8$ and yields only $+2.8$ in fast-thinking evaluation.
Chat is not available.
Successful Page Load