LoRASpace: A Pool-Wide Shared Substrate for Static and Dynamic Multi-LoRA Composition
Abstract
Low-rank adaptation (LoRA) attaches small low-rank matrices to a frozen backbone, and public repositories like Hugging Face host many such LoRAs per backbone, each for a different concept, character, style, or task. Users routinely want to combine several LoRAs in a single generation, but two existing strategies involve a trade-off with each other. First, combining the LoRA weights into a single update before inference (e.g., static merge) needs only one forward pass, but quality drops, because LoRAs trained for different concepts interfere when their weights are summed into a single update, distorting each LoRA's contribution. Second, running each LoRA in a separate forward pass and summing the outputs (e.g., composite) preserves output quality, but each LoRA adds a full backbone pass, so inference time scales linearly with the active LoRA count. To recover composite-level quality at merge cost, we study a property of LoRA pools overlooked by prior multi-LoRA work. When LoRAs in a pool are trained on related concepts, their low-rank factors overlap substantially. Using singular value decomposition, we find the minimal subspace that preserves 99% of the factor energy. On three diffusion pools, only 46% to 57% of the sum of individual LoRA ranks is enough to represent the pool, and the subspace grows sub-linearly with pool size. A 48-LoRA sub-pool of LoRA-Hub on FLAN-T5 NLP tasks shows comparable overlap (about 62%), while the diversity-curated LoRA-Retriever benchmark sits near the no-overlap baseline at 96%, indicating that across these five pools pool composition rather than modality predicts where overlap is large. We use this overlap to construct LoRASpace, a training-free pool-wide method where the subspace matrices are computed once per pool and reused across two inference modes: a static mode at merge inference cost, and a dynamic mode for input-dependent routing at substantially lower cost than composite. On the diffusion pools, static LoRASpace matches composite quality at merge cost, and dynamic LoRASpace achieves more than 6× speedup over composite at six active LoRAs on a 17-LoRA pool. In our ablation, we show that the quality gain comes from truncating to this low-rank shared subspace, which acts as a form of regularisation. On the near-orthogonal LoRA-Retriever benchmark, the framework predicts limited compression benefit, with gains coming from external retrieval coupling.