What Should Remain After Forgetting? Rethinking LLM Unlearning as Predictive Posterior Correction
Abstract
LLM unlearning is often implemented by assigning surrogate targets, such as refusal, uniformity, or likelihood suppression, to forget-related prompts. We argue that this surrogate assignment view is misaligned with the counterfactual goal of unlearning: recovering the retain-only predictive posterior. This mismatch is especially problematic in mixed-evidence regimes, where a response may be associated with forgotten evidence while still being partially supported by retained evidence. We propose EASE: Evidence Attribution and Subtraction Estimator, a logit-space posterior-correction method that estimates the forget-induced component of the deployed model's predictive support and subtracts it at inference time. EASE uses lightweight deletion and compensation assistants to remove forget-neighbour evidence while restoring nearby retain-supported evidence. We formalize the suppression-scope dilemma and provide guarantees showing when posterior correction recovers the retain-only posterior. Experiments on TOFU and MUSE show that EASE improves forgetting--retention trade-offs over strong baselines, achieving the best aggregate score in 5/6 main TOFU settings, reducing mixed-query semantic leakage by 22.2%, and matching or exceeding full-retain baselines with 9 times fewer retain examples. Our code is available at: https://anonymous.4open.science/r/EASE-9675.