Phase-MoE: A Co-Design to Bound the Expert Explosion in Block Diffusion Language Models
Abstract
Block Diffusion Language Models (BDLMs) have emerged as a compelling paradigm bridging autoregressive modeling and discrete diffusion, pairing inter-block causal generation with intra-block bidirectional parallel decoding. To scale these architectures efficiently, integrating Mixture-of-Experts (MoE) is essential. However, combining MoE with block-parallel decoding triggers a systemic Expert Explosion: processing multiple tokens simultaneously causes their independently routed experts to rapidly span the Expert pool. This severely inflates memory traffic from High-Bandwidth Memory (HBM) to SRAM, creating a memory-bound bottleneck. Existing post-hoc mitigations---such as sequence-level sharing, dynamic caching, or speculative exploration---face scaling limitations under production continuous batching constraints, where the combined temporal and semantic heterogeneity of asynchronous requests demands disjoint expert subsets, escalating global memory traffic. To resolve this, we propose a bandwidth-aware co-design comprising Phase-Constrained MoE and Phase-Aware Scheduler. Recognizing that discrete diffusion progresses through predictable mask-density phases, we explicitly condition the router's candidate expert pools on the temporal denoising phase during training, leaving compute-bound prefill unconstrained. The Phase-Aware scheduler subsequently clusters asynchronous requests by mask density at runtime, guaranteeing overlapping expert execution. Under production-scale continuous batching simulations, our co-design reduces unique expert activations by up to 56.2\% and cuts MoE kernel latency by up to 38.5\% compared to standard serving baselines, preserving competitive reasoning and generative accuracy. Anonymized code repository: \url{https://anonymous.4open.science/r/dllm_project-7B77}