Auditing Post-Training Unlearning in Diffusion Language Models using Sparse Autoencoders
Yunha Seo ⋅ Changhoon Kim
Abstract
Machine unlearning is an approach to privacy protection: it aims to remove the influence of desingated data while preserving the model's behavior on information that should be retained. While growing body of work investigates objective in autoregressive large language models, machine unlearning remains underexplored in diffusion language models (DLMs). DLMs generate text through iterative denoising, exposing partially maasked states their internal representations change throughout generation process. This raiss a DLM-specific question: can selective forgetting improve by identifying forget-data-associated features and adapting their suppression across denoising stages? We conduct a controlled investigation of this question in LLaDA-8B-Instruct. To evaluate the unlearning of knowledge plausibly encoded during pre-training, we use WMDP as our forget-domain benchmark. Using top-$k$ sparse autoencoders and denoising-state replay, we compare static and stage-adaptive feature suppression strenghts. We find substantial shifts in forget-data-associated features across denoising stages and demonstrate that traking these dynamics alone does not reliably improve the forget-retain trade-off. Together, our study provides an initial empirical foundation for developing and evaluating machine unlearning as a post-training privacy mechaniss in DLMs.
Chat is not available.
Successful Page Load