LLM Routing Through the Lens of Recommendation: A Roadmap for Efficient AI Orchestration
Abstract
The rapid proliferation of Large Language Models (LLMs) has created a diverse ecosystem and motivated model routing as a way to navigate differences in architectures, capabilities, latency, and cost. Yet current routing methods still face bottlenecks in generalization, scalability, and interpretability. To solve these problems, we argue that model routing can be formulated as a recommendation problem. This formulation links routing bottlenecks to problems long studied in recommendation systems (RecSys), including cold start, large-scale recommendation, dynamic adaptation, and explainable decision-making. Building on this RecSys perspective, we organize a roadmap that maps major routing bottlenecks to established RecSys paradigms, providing a practical foundation for adapting mature recommendation methodologies to multi-LLM orchestration. We further present preliminary empirical case studies showing that recommendation-inspired routers can achieve strong accuracy-cost trade-offs while supporting model cold-start and interpretable routing behavior.