HAPS: Hierarchical LLM Routing with Joint Architecture and Parameter Search
Abstract
Large language model (LLM) routing aims to exploit the specialized strengths of different LLMs for diverse tasks. However, existing approaches typically focus only on selecting LLM architectures, treating candidate models as static black boxes while overlooking parameter adaptation, which can substantially affect task performance. In this paper, we introduce HAPS, a hierarchical LLM routing framework that jointly searches over model architectures and parameters. Specifically, a high-level router selects among candidate LLM architectures, while a low-level router generates input-conditioned LoRA parameters for the selected architecture. To couple these two decisions, we design a shared parameter generation mechanism that enables cross-level knowledge transfer between architecture routing and parameter adaptation. We further optimize the whole framework with a reward-augmented training objective. Extensive experiments show that HAPS consistently outperforms strong routing baselines. We further conduct broad ablations and analyses on parameter sharing, architectural choices, scalability, and efficiency, providing additional evidence for the effectiveness and practicality of the proposed framework. We have released our code at https://anonymous.4open.science/r/HAPS_private-68CB.