Chain-of-Route: State-Aware LLM Routing for Multi-Turn Conversations
Abstract
LLM routing dispatches each user query to the most suitable model from a pool of candidates, balancing response quality against serving cost. As LLMs are increasingly deployed in conversational interfaces, routing decisions must now be made repeatedly within an ongoing session. Existing routers treat each query independently, implicitly assuming that the optimal model depends only on the current input. This assumption ignores two forms of session state that accumulate across turns: the dialogue state, which shapes query meaning and difficulty, and the \textit{serving state}, which determines each model's true \textit{serving cost} via KV-cache reuse. Neglecting dialogue state leads to misrouting when queries depend on prior turns, while ignoring serving state causes systematic cost overestimation that worsens as conversations grow longer. A practical multi-turn router must therefore capture dialogue state for accurate quality prediction, and track per-model serving state for accurate cost estimation. To address this gap, we introduce Chain-of-Route (CoR), the first state-aware routing framework for multi-turn conversations. CoR comprises two modules: a Dialogue State Propagator (DSP) that maintains a compact recurrent state to capture session-level context across turns, and a Serving State Tracker (SST) that tracks per-model KV-cache to compute true serving costs. Together with the framework, we construct a unified multi-turn routing benchmark from WildChat and LMSYS-Chat-1M with per-turn quality annotations across five LLMs. Experiments show that CoR achieves state-of-the-art quality-cost tradeoffs, reducing QNC by up to 20.3\% over the strongest baseline while improving response quality.