Holistic Scaling Laws for Optimal Mixture-of-Experts Architecture Optimization
Weilin Wan ⋅ Jingtao Han ⋅ Debing Zhang ⋅ Weizhong Zhang ⋅ Cheng Jin
Abstract
Scaling laws govern macroscopic resource allocation for LLMs, yet precise architectural configurations for Mixture-of-Experts (MoE) models remain guided by heuristics. Existing MoE scaling studies either incorporate MoE-specific variables into scaling formulas, causing fitted coefficients to grow rapidly without proportional experimental support, or fix all non-MoE factors, implicitly assuming global architecture does not influence local MoE scaling. We propose a reusable framework for holistic MoE architectural optimization. We first reveal that relying solely on FLOPs per token ($M$) biases MoE evaluation, as heterogeneous Attention/FFN computational densities enable parameter inflation without effective compute gains. We therefore establish a joint constraint triad of $M$, active parameters ($N_a$), and total parameters ($N$) for rigorous MoE characterization. To tame the resulting $\mathcal{O}(n^{16})$ search space, we employ mathematical decoupling: structural constraints and a rank-preserving property of the hidden dimension factorize the optimization into an $\mathcal{O}(n^3) + \mathcal{O}(n^2)$ two-phase search. Through extensive validation across 670+ MoE models spanning $10^{18}$ to $3 \times 10^{20}$ FLOPs, we derive globally applicable scaling laws mapping any compute budget to its optimal architecture. A key finding is that the near-optimal configuration band widens with scale, providing a quantitative basis for trade-offs between scaling law recommendations and engineering constraints. Our work delivers actionable blueprints for optimal MoE design under arbitrary compute budgets.
Chat is not available.
Successful Page Load