DivMoE: Fine-Grained MoE Upcycling via Cross-Domain Expert Composition
Abstract
Mixture-of-Experts (MoE) architectures have become essential for building capable large language models, with recent work demonstrating the benefits of fine-grained expert designs. However, training MoE models from scratch is computationally expensive, and existing upcycling methods that convert dense models to MoE face limitations: they either initialize experts as identical copies (limiting routing diversity) or use random perturbations (risking knowledge loss). We propose \textsc{DivMoE}, a framework that addresses these challenges through two innovations. First, we introduce \emph{domain-specialized expert initialization}, deriving fine-grained experts from models fine-tuned on distinct domains (mathematics, code, science, commonsense), providing meaningful diversity while preserving pre-trained knowledge. Second, we propose \emph{diversity-constrained routing}, which enforces that each token selects at most one expert per domain group, structurally preventing routing collapse and enabling cross-domain knowledge composition. Experiments on two base models demonstrate that \textsc{DivMoE} outperforms all upcycling baselines (47.0% vs.\ 44.4% average accuracy) and achieves competitive performance with MoE models trained from scratch, matching Moonlight-MoE at 64.5% average while being constructed via efficient upcycling.