FELPS: Fair and Efficient Scheduling for Multi-LoRA Serving System
Abstract
Low-Rank Adaptation (LoRA) is a popular parameter-efficient fine-tuning method for Large Language Models (LLMs), and practical systems usually need to serve many concurrent LoRA adapters to process different tasks. This is challenging due to the inherent tension between \textit{fairness} and \textit{efficiency}, i.e., achieving fairness requires frequent switching among the adapters such that they have similar service quality, while good efficiency requires to minimize adapter switching to reduce memory movement overheads. Existing schedulers can only achieve a fixed trade-off between fairness and efficiency, and their trade-offs are often suboptimal. To tackle this problem, we propose \textbf{F}airness-aware \textbf{E}fficient \textbf{L}oRA \textbf{P}riority \textbf{S}cheduling (FELPS), which explicitly decomposes fairness and efficiency considerations. In particular, for fairness, FELPS adapts the classic proportional fairness and ensures that the amount of service received by each adapter is proportional to its workload. For efficiency, FELPS prioritizes adapters that are already loaded to GPU memory. To conduct scheduling, FELPS combines to the fairness and efficiency terms with an adjustable weight to achieve arbitrary trade-offs between fairness and efficiency. Evaluations on real workloads show that FELPS achieves a superior trade-off between fairness and efficiency than all baselines.