CTRL: Continual Test-Time Reinforcement Learning for Large Language Models
Abstract
Test-Time Reinforcement Learning (TTRL) constructs pseudo-supervisory signals by sampling reasoning trajectories during inference, enabling online adaptation to current inputs. However, existing research predominantly focuses on single-task settings, overlooking continuous updates across task streams. Theoretical and empirical analyses reveal two coupled challenges in this setting: (1) Error Accumulation: Majority voting tends to treat erroneous consensus as pseudo-labels, with such bias amplifying across sequential parameter updates; (2) Catastrophic Forgetting: Gradients from new tasks conflict with historical knowledge, causing acquired reasoning patterns to be overwritten. These issues form a positive feedback loop: pseudo-label noise exacerbates gradient interference, while degradation of historical knowledge reduces subsequent trajectory quality, jointly driving performance deterioration during continual adaptation. To address this, we propose the CTRL framework: first, we introduce a Process Reward Model (PRM) to replace outcome-level voting, filtering high-confidence trajectories via fine-grained process scoring to suppress pseudo-label bias; second, we design a cognitive anchor-driven soft gradient correction module that dynamically recomputes anchor gradients from a historical trajectory memory pool and projects out update components negatively correlated with historical knowledge, constraining the parameter update trajectory. Experiments demonstrate that CTRL significantly outperforms existing TTRL baselines on continuous task streams, exhibiting stronger robustness in blocking error propagation and preserving historical reasoning patterns.