Hallucination-Guided Unlearning: Using Hallucination Traces to Reveal Overfitted Memories
Abstract
Selective unlearning in large language models is increasingly necessary for safety and compliance, yet harmful or memorized content is often latent, entangled, and hard to specify as a clean ``forget set.'' We aim to close this gap by identifying what to forget from the model’s own failure modes rather than relying solely on explicit supervision. We propose Hallucination-Guided Unlearning (HGU), which treats hallucinations as diagnostic signals: it localizes unsupported spans via evidence and internal consensus, traces a sparse set of high-responsibility MLP neurons, and edits them with compact gates under a utility-preserving KL constraint. Across TOFU, HGU achieves the best MU--FE trade-off with Avg of 0.7221 (forget05) and 0.7699 (forget10), and it also attains the top SepS H-Avg (0.4382/0.4715) under mixed forget/retain prompts; on WMDP auxiliary, it yields the strongest forgetting (e.g., 14.2 on Economics) while improving Retain (52.7) and MMLU (57.2). These results suggest that reframing hallucinations as actionable traces enables targeted, post-hoc unlearning that better balances forgetting efficacy and general utility without full retraining.