METAFORGET: Audit-Driven Update-Policy Learning for Reliable Language Model Unlearning
Abstract
Deployed language models are increasingly expected to remove private, copyrighted, or hazardous content without full retraining. Existing unlearning methods remain brittle because they usually optimize a fixed forgetting objective with a global update schedule: weak updates leave paraphrased knowledge intact, while stronger updates can damage retained capabilities or nearby facts. We identify this failure mode as update homogeneity. To address it, we introduce METAFORGET, a meta-optimization framework that treats unlearning as learning an audit-driven update policy rather than designing another standalone forgetting loss. A lightweight controller maps memorization, paraphrase-sensitivity, and signed forget-retain conflict signals to span weights, layer gates, step sizes, and retain-gradient projection strengths. The controller is trained with a bilevel objective that differentiates through short LoRA unlearning trajectories and evaluates the resulting model on held-out forget, retain, locality, and leakage audits. This policy can instantiate NPO- or RMU-style losses while changing where and how strongly they act. Across TOFU, WMDP, and MUSE-style sequential deletion settings, METAFORGET improves the forget-retain-robustness frontier over strong baselines, moves loss-based membership-inference diagnostics closer to chance, and better preserves utility under repeated requests. The results support a view of practical LLM unlearning as audit-driven policy learning: the central object is not only the forget loss, but the localized update that must pass the deletion audit.