Reconstruction-Optimized Expert Pruning for Mixture-of-Experts Language Models
Abstract
Mixture-of-Experts (MoE) models reduce computation through sparse expert activation, but deploying them still requires storing a large number of expert parameters. We study post-training expert pruning using task-domain calibration data as a subset-selection problem: given a pretrained MoE layer, select a fixed number of experts that best reconstruct its original output. Exact subset search is combinatorial and becomes intractable for modern MoE models with hundreds of experts per layer. We develop a scalable continuous relaxation that directly optimizes differentiable expert selection variables, followed by a lightweight reconstruction-based update of the retained expert weights. Across Qwen3-30B-A3B, Qwen3-235B-A22B, and Mixtral-8x7B, our approach consistently improves over existing expert pruning methods, with gains of up to 66.0% on MATH-500 and 39.5% on LiveCodeBench for Qwen3, and up to 13.7% on GSM8K and 12.7% on MBPP for Mixtral. The improvements are particularly pronounced under aggressive pruning regimes.