Don't Skip the Baseline: A Controlled Comparison of Dynamic Expert Skipping in MoEs
Jakob Steimle ⋅ Thomas Elsken ⋅ Daniel Ehrhardt ⋅ Lukas Rinder ⋅ Daniel Cremers ⋅ Xi Wang
Abstract
A growing body of work proposes dynamic expert-skipping methods for Mixture-of-Experts (MoE) models that adjust the number of experts activated per token during inference, aiming to preserve model quality while further reducing the number of active parameters. These methods have added calibrated importance scores, inter-expert similarity metrics, or learned per-token allocators in order to identify experts to skip. However, different calibration regimes, different reductions in the effective number of experts, and evaluation suites make comparisons difficult. We classify four representative prior methods into a unified taxonomy and benchmark them against a natural, simple baseline: statically reducing the number of experts $k$ (i.e., same reduced $k$ across all layers and tokens). We provide a fair comparison of these methods at matched expert-compute budget, on two modern MoEs (Gemma4-26B-A4B and Qwen3-30B-A3B) in a text-only setting and a contemporary vision-language model (Qwen3-VL) in a multimodal setting. Surprisingly, we find that the simple baseline is hard to outperform consistently across models, modalities, benchmarks, and calibration settings. The methods we benchmark often underperform the baseline and achieved gains, when present, are minor. We further find that the effective $k$ measured (or optimized for) in calibration does not necessarily match the effective $k$ in testing, resulting in over- or under-estimating the compute at inference time. Based on these findings, we recommend that practitioners use simple, calibration-free static $k$ reduction, split by modality for multimodal models and validate this recommendation on a held-out model and modality.
Chat is not available.
Successful Page Load