DART: Zero-Shot Dual-Side Alignment Routing for LLM Performance-Cost Tradeoffs
Abstract
As the ecosystem of Large Language Models (LLMs) rapidly expands, LLM routing has become essential for dynamically balancing performance and cost. However, existing methods fail to simultaneously satisfy the three core criteria of an ideal router, i.e., high routing precision, minimal intrinsic operational overhead, and zero-shot onboarding of new models. These methods rely heavily on surface-level semantics and discrete ID mappings, conflating mere semantic similarity with underlying capability requirements and thereby precluding zero-shot, training-free generalization. To address these limitations, we propose DART, a novel zero-shot routing framework grounded in dual-side alignment to optimize LLM performance-cost tradeoffs. Instead of relying on direct semantic matching, we decouple the routing process into a demand-supply alignment problem within a shared low-rank latent capability space. Specifically, on the supply side, DART constructs latent capability embeddings for candidate models, using an initialization mechanism to anchor newly introduced models in this latent space via metadata and family-graph prior, without retraining. On the demand side, a dedicated encoder explicitly distills the query's directional capability requirements and inherent difficulty into a latent demand embedding. This decoupled representation allows routing decisions to be computed via a lightweight dot-product operation between the demand and capability embeddings, modulated by query-adaptive cost sensitivity, minimizing overhead. Extensive experiments across comprehensive benchmarks demonstrate that DART achieves state-of-the-art utility under diverse cost constraints, satisfying the three core criteria with exceptional robustness in cold-start scenarios.