Distributionally Robust Mixture-of-Experts Training
Xin Teng ⋅ Muxiao Li ⋅ Hongyi Wen
Abstract
Mixture-of-Experts (MoE) transformers scale capacity by activating only a few experts per token, but this sparsity creates a hidden reliability problem: when routing is imperfect, load-balanced models may send tokens to experts that are insufficiently trained for the assigned inputs. We propose Distributionally Robust MoE Training (DRMoET), a drop-in objective that treats layer-wise experts as endogenous robustness groups and optimizes high-loss routing outcomes rather than merely equalizing traffic. DRMoET updates a per-layer expert distribution by an entropy-regularized softmax rule on EMA-smoothed, activation-weighted expert losses, strengthening plausible non-top routing paths while preserving standard MoE computation. Under the FLAME-MoE recipe at 746M-total and 10.3B-total scales, DRMoET improves downstream averages over both standard FLAME-MoE and auxiliary-loss-free balancing. Mechanistic analyses show lower expert-loss variance with nearly unchanged mean loss, 4.3\% lower degradation under forced mid-$k$ misrouting, and improved domain--expert specialization. These results position routing robustness--not only utilization balance--as a practical objective for reliable sparse MoE scaling.
Chat is not available.
Successful Page Load