From Anti-Forgetting to Fast Adaptation: Continual Reinforcement Learning with World Models
Qiyang Zhou ⋅ Yuliang Cai ⋅ Peng Wang ⋅ Senquan Yi ⋅ Yanming Li ⋅ Haobo Fu ⋅ Li Shen
Abstract
Continual Reinforcement Learning (CRL) requires agents to adapt to sequentially arriving tasks while retaining performance on previous ones. Most existing CRL methods rely on model-free frameworks, addressing catastrophic forgetting via regularization, parameter isolation, or experience replay. World models offer a promising alternative by accumulating transferable dynamics knowledge and enabling planning-based decision-making. However, existing world model approaches suffer from slow adaptation at task boundaries: abrupt reward shifts mislead the planner, fixed planning horizons accumulate prediction errors when the dynamics model is unreliable, and quality-agnostic replay dilutes training while exacerbating forgetting. We propose HERD (Horizon-adaptive Elite Replay with reward Disagreement), which addresses these challenges from two complementary perspectives: for $\textbf{\emph{how to learn}}$, Uncertainty-aware Reward Planning leverages ensemble disagreement to accelerate reward calibration and guide exploration, while Adaptive Planning Horizon adjusts rollout depth based on dynamics prediction error to prevent error accumulation; for $\textbf{\emph{what to learn}}$, Elite Experience Replay applies quality-stratified reweighting of historical trajectories to strengthen knowledge retention. Together, these components $\textbf{\emph{accelerate new-task adaptation and mitigate catastrophic forgetting.}}$ Experiments on Continual World, Continual Bench, and DM Control show that HERD consistently outperforms existing methods in overall performance, forgetting mitigation, and new-task adaptation.
Chat is not available.
Successful Page Load