Machine Unlearning in Diffusion LLMs
Abstract
Machine unlearning (MU) aims to remove sensitive or undesired knowledge from a trained model without retraining from scratch. While MU has been widely studied for autoregressive large language models, its application to diffusion large language models (DLLMs) remains largely unexplored. In this paper, we present the first systematic study of MU in DLLMs and show that existing unlearning objectives transfer poorly to this setting due to sparse supervision and inaccurate forgetting localization. To address these challenges, we propose \textsc{DL-Eraser}, a DLLM-native unlearning framework that suppresses target recoverability under high-mask conditioning while constraining updates to a utility-preserving subspace. Extensive experiments show that \textsc{DL-Eraser} achieves stronger forgetting while better preserving model utility than existing baselines. To the best of our knowledge, this is the first work that systematically investigates MU in DLLMs.