Evaluating Deployable Inference-Time Error Prediction in Vision MoEs
Abstract
Sparse mixture-of-experts (MoE) models raise a natural inference-time question: can experts not selected by the router improve predictions or estimate their errors without retraining? We study this question in vision MoEs across five RGB and multispectral classification datasets and multiple expert counts. We evaluate additional expert inference for three outcomes: top-1 accuracy, raw ranking of default router errors, and calibrated prediction of whether the router-selected classification is wrong. An upper-bound analysis shows that non-selected experts often contain useful class signal when the routed path fails. However, this signal does not reliably improve deployable top-1 accuracy, and raw multi-forward uncertainty scores yield only limited gains for ranking router errors. The strongest benefit appears in calibrated router-error prediction: post-hoc models that combine router signals with features from non-selected experts reduce Router-error Expected Calibration Error (ECE) by roughly 10-13% on average, with several settings exceeding 20% relative improvement over a calibrated router-only baseline. Overall, additional inference-time expert usage in vanilla vision MoEs is most useful for estimating the reliability of the routed decision, rather than directly improving classification accuracy.