CLUE: Closing the Loop on Conflict and Collapse in LLM Unlearning
Abstract
Machine unlearning for large language models optimizes two opposing losses: a forget objective drives the model away from targeted knowledge, while a retain objective holds it near its original behavior. Existing methods commit to a fixed schedule of learning rate, retain weight, and step count chosen by offline grid search, which is blind to how the trajectory actually evolves. We show that LLM unlearning trajectories share a method-agnostic three-phase structure, an early phase where the two gradients are nearly orthogonal, a conflict phase where they oppose each other and retain loss begins to grow, and a terminal phase where representational divergence crosses an irreversibility threshold. Each phase calls for a different control policy. We propose CLUE, a closed-loop framework whose architectural choices each correspond to a phase or transition: a dual-tower controller separates forget and retain signal pathways, a conflict-driven gate switches between them, and an explicit collapse-risk head supervises anticipatory stopping. The controller is trained by task-distribution meta-optimization. We provide a single-step bound on conflict-induced retain loss growth, a quantitative irreversibility result, and a stopping regret bound. On TOFU, WMDP, and MUSE with Llama-2/3 and Mistral models from 7B to 70B, CLUE achieves better forget-quality versus model-utility trade-offs than fixed-strategy baselines, stops within 5\% of the oracle, and remains robust under adversarial prompting, relearning, and INT4 quantization.