When LLM Routers Overpay: Strong-Model Over-Selection under Loose Budgets
Abstract
LLM routing is designed to avoid overpaying for inference: easy queries should be handled by cheaper models, while stronger models should be reserved for cases where they are truly needed. However, we show that existing routers often violate this principle under loose cost budgets. As the budget increases, they increasingly route queries to the strongest and most expensive model, even when cheaper candidates achieve comparable or identical outcomes. We refer to this behavior as strong-model over-selection under loose budgets. Across both unimodal and multimodal routing benchmarks, this behavior weakens the cost-saving motivation of routing without necessarily improving quality. We diagnose this phenomenon through the mismatch between common router training objectives and deployment-time routing decisions. Many routers are trained to predict scalar performance scores, whereas cost-aware routing ultimately depends on query-specific comparisons among budget-feasible models, with cost used to distinguish equivalent or near-equivalent choices. In small-margin regimes, modest prediction errors can therefore flip relative orderings and induce cost-insensitive selections. Motivated by this diagnosis, we propose EquiRouter, a decision-aligned router that directly learns query-dependent model rankings with a Cost-Aware Ranking Objective and a lightweight Query--Model Interaction Representation. Experiments across RouterBench, MMR-Bench, MixInstruct, and RouterEval show that EquiRouter reaches the strongest model-level performance with lower relative cost than compared baselines, demonstrating consistent improvements across text-only, multimodal, continuous-utility, and large-model-pool routing settings.