Mask-Conditioned Gradient Masking for Fine-Tuning Mixture-of-Experts Diffusion Language Models
Abstract
While autoregressive models dominate the LLM landscape, Discrete Diffusion Language Models have emerged as a compelling alternative due to their advantages in parallel decoding and bidirectional contextual modeling. To enhance scalability and efficiency, recent works have integrated the Mixture-of-Experts architecture into the DLM framework. However, MoE-DLMs remain underexplored. In this work, we present an empirical study revealing that MoE-DLMs exhibit not only task-level expert specialization, but also exhibit specialization across different mask-rate diffusion regimes. We further show that during fine-tuning, gradients induced by mismatched mask rates can interfere with the updates of regime-specialized experts, leading to suboptimal adaptation. Motivated by this, we propose MaGM, an adaptive mask-conditioned gradient masking method for MoE-DLM fine-tuning, which dynamically masks expert parameter updates based on the mask rate of training instance. Experiments demonstrate that MaGM consistently outperforms standard full-parameter fine-tuning for diffusion language models, validating the benefits of regime-aware expert adaptation. Our code will be publicly released.