What Do Expert Interventions Measure in Mixture-of-Experts Models?
Sabbir Ahmed ⋅ Ahnaf Adib ⋅ Pritom K Saha ⋅ Latifur R Khan
Abstract
Mixture-of-Experts (MoE) interpretability studies characterize experts, remove them to locate behaviors, and rescale them to steer behavior. We compare routed experts with active neurons as intervention units for individual predictions. Direct attribution to the output token selects neurons, and the same multiplicative operation scales either those coordinates or routed experts. We evaluate Qwen3-30B-A3B base and post-trained models, Gemma-4-26B-A4B, Mixtral-8x7B on factual recall, Wikipedia text, arithmetic, multi-step reasoning, hallucination, multilingual, sentiment and safety tasks. Deactivating a small attribution-selected neuron set reduces target-token probability more than deactivating routed experts, whose removal can instead raise it. Neuron amplification helps only when the target begins at low probability; for already favored targets, the effect is negligible or negative, while expert suppression can be up to 9.5$\times$ larger than amplification. At the final layer, expert-deactivation magnitude follows removed routing mass more closely than expert selection by routing weight. Transcoder features reproduce the separation between attribution-selected and random deletions, but amplification averages near zero across 172 features.
Chat is not available.
Successful Page Load