Confidence Estimation via Decoupled Smoothing for Dynamic LLM Routing and Aggregation
Abstract
The rapid emergence of heterogeneous large language models (LLMs) has highlighted the limitations of static single-model inference systems. Although model routing alleviates this issue by assigning the most suitable model to each query, existing methods often struggle to stably align query semantics with actual model performance and remain constrained by the capability ceiling of a single model. On the other hand, although multi-model aggregation can improve performance on complex queries, it also introduces extremely high computational overhead for routine tasks. In this paper, we propose PARS, a model confidence estimation method for dynamic LLM routing and aggregation. PARS first decouples the feature representation space and applies structured smoothing to local performance signals, producing more discriminative confidence score estimates for candidate models. On this basis, PARS represents inference modes such as direct routing, self-consistency sampling, and multi-model voting as test-time compute allocation actions under a budget constraint. By extracting state features such as routing confidence distributions and predictive uncertainty, a lightweight selector adaptively assigns an inference strategy to each query. Extensive experiments show that PARS improves average accuracy over the best baselines by 1.02% and 1.36% in the single-model and multi-model settings, respectively, while reducing average inference cost by 1.73x and exhibiting stable out-of-distribution generalization.