Learning What to Forget: Improving LLM Unlearning via Learned Token-Level Importance
Abstract
Machine unlearning aims to remove targeted knowledge from a trained model while preserving its general capabilities. Doing so effectively for auto-regressive language models requires identifying which tokens in a forget sample are truly relevant to forgetting, as uniformly applying unlearning across all tokens can degrade utility. Existing approaches, however, either ignore this heterogeneity or rely on auxiliary models, hand-crafted heuristics, or external annotations to approximate token-level relevance. We instead characterize token-level forget relevance through its interaction with the retain objective: tokens are forget-relevant to the extent that minimizing the forget loss on them does not conflict with the retain objective. Based on this insight, we introduce Alternating Token-Weighted Unlearning (ATWU), a framework that jointly learns token relevance and model parameters during the unlearning process. ATWU uses a lightweight linear scorer over model hidden states to predict forget-relevance scores with negligible computational overhead and no external supervision. Experiments on TOFU and RWKU show that ATWU achieves state-of-the-art forget--retain trade-offs, outperforming sample-level methods, probability-based token-weighting heuristics, and auxiliary-model-based approaches. Moreover, the learned scorer aligns with ground-truth forget-relevant spans substantially better than existing methods, suggesting that ATWU learns semantically meaningful token-level forgetting signals. Overall, ATWU shows that token-level forget-relevance can be effectively inferred from model representations during unlearning, enabling efficient and selective forgetting.