ExpertNavigator: Functionally Coherent Expert Grouping and Pairwise-Ranked Routing for High-Fidelity Dense-to-MoE Conversion
Abstract
Dense large language models (LLMs) exhibit implicit sparsity at inference time: for a given token, only a small subset of FFN neurons contributes dominantly. This observation motivates Dense-to-MoE conversion, which partitions FFN neurons into experts and employs a lightweight router to sparsely activate experts for efficiency. However, existing conversion methods often incur non-trivial quality degradation under a fixed activation budget, largely due to suboptimal expert construction and unstable router learning in fine-grained settings. Motivated by this, we present ExpertNavigator, which improves Dense-to-MoE conversion with a more functionally coherent expert grouping method and a more stable and effective pairwise-ranked router training strategy, leading to markedly better quality under the same activation budget. Moreover, we further introduce an adaptive layer-wise sparsity allocation strategy to better utilize a global activation budget, and apply STE-based joint training to enhance model--router compatibility under hard routing. Experiments across various benchmarks show that ExpertNavigator consistently outperforms prior Dense-to-MoE approaches under different sparsity budgets, retaining up to 97.8% of the dense model's average score while activating only 60% of FFN parameters, substantially narrowing the gap to dense models and delivering practical efficiency gains.