Token-Conditional Expert Dropout: Implicit Regularization for Stable MoE Pretraining
Abstract
Sparse Mixture-of-Experts (MoE) layers underpin the most capable open-weight large language models, including DeepSeek-V3.2, Qwen3-MoE, and the Llama 4 family, but their pretraining is plagued by an early-stage pathology in which a handful of experts dominate the routing distribution while the remainder are starved of gradient signal. Existing remedies treat the symptom rather than the cause: load-balancing auxiliary losses, sequence-level balance constraints, and router z-loss penalties all push the gating distribution toward uniformity through additional gradient signals that compete with the language-modeling objective and require careful coefficient tuning. We argue that the root issue is a positivefeedback loop between expert capability and expert selection, and we propose Token-Conditional Expert Dropout (TCED), a training-time perturbation that masks the chosen expert for a token with probability proportional to that expert’s recent utilization and re-routes the token to the highest-scoring alternative. TCED is parameter-free, requires no auxiliary loss, and provides implicit regularization analogous to attention dropout while breaking the collapse feedback loop. We derive an unbiased gradient estimator under TCED, prove a variance-reduction lemma that bounds the variance of normalized expert utilization, and pretrain a 1.4B-active / 8B-total MoE on 480B tokens of a high-quality web mixture. TCED reduces expert-utilization variance by 67% across training, lifts MMLU-Pro by 2.1 points and GPQA by 2.1 points relative to a vanilla MoE, and matches a heavily tuned auxiliary-loss baseline without a single coefficient sweep. Results are preliminary single-seed estimates; final-version revisions will report multi-seed confidence intervals.