CURe: Conservative Unlearning with Soft-Gating Regularization for Offline Reinforcement Learning
Abstract
Offline reinforcement learning (RL) trains policies from static datasets, but deployed agents may later need to attenuate the influence of specific data for privacy, safety, or fairness. This is challenging in the overlap regime, where the forget set and retained data share state-action support: unlearning cannot be uniform as suppressing values on shared support can inadvertently erase behaviors required for high performance. We propose Conservative Unlearning with soft-gating Regularization (CURE), a selective unlearning framework tailored to overlap. CURE assigns each forget trajectory an overlap-interference score via cosine alignment between its temporal difference (TD) semi-gradient and a retained reference gradient, and uses it to soft-gate i) conservative, state-conditional critic suppression relative to a retained baseline and ii) coupled actor updates, selectively modulating forgetting based on interference, reducing impact on aligned data. We provide a first-order analysis showing that CURE descends a forget surrogate while controlling retained drift via alignment-dependent coupling. Across MuJoCo benchmarks and offline RL backbones, CURE consistently improves the forget-utility trade-off over baselines, achieving low auditor-detectable forget influence with near-original returns at a fraction of retraining cost.