MoE Routing Testbed: Studying Expert Specialization and Routing Behavior at Small Scale
Abstract
Sparse Mixture-of-Experts (MoE) architectures are increasingly popular for frontier large language models (LLMs) but they introduce training challenges due to routing complexity. Fully leveraging parameters of an MoE model requires all experts to be well-trained and to specialize in non-redundant ways. Assessing this, however, is complicated due to lack of established metrics for specialization. We propose the MoE Routing Testbed, a setup that gives clearer visibility into routing dynamics at small scale while using realistic data. The testbed pairs a data mix with distinguishable domains with a reference routing based on these domains. This enables quantifiable measurement of expert specialization. To demonstrate the value of the testbed, we compare various MoE routing approaches and can clearly show that balancing scope is the crucial factor that allows specialization while maintaining high expert utilization.