Adaptive Fine-Tuning Scheduler for Multi-Tenant Edge LLM via Convergence-Aware Bandits
Yandi Li ⋅ Jianxiong Guo ⋅ Yupeng Li ⋅ Zhiqing Tang ⋅ Tian Wang ⋅ Weijia Jia
Abstract
Personalizing Large Language Models (LLMs) directly on local edge servers is becoming increasingly important for privacy-preserving and context-aware applications. However, this potential is bottlenecked by hardware resources: limited GPU memory permits only a fixed number of train-ready LoRA adapters, while compute constraints enforce sequential fine-tuning updates. This creates a critical challenge: scheduling scarce update opportunities across a dynamic stream of resident tenants to improve overall model quality and scheduling stability. Crucially, standard Multi-Armed Bandit (MAB) algorithms fail to distinguish between low-potential tenant streams and fully saturated adapters, leading to wasteful exploration on tasks that offer little further gain. To address this, we propose WCA-UCB (Windowed Convergence-Aware UCB), a scheduler designed for this piecewise-converging non-stationary environment. By formally modeling fine-tuning as a slot-constrained bandit problem with piecewise-converging rewards, WCA-UCB detects when an adapter has saturated to pause its training and resets stale statistics after tenant replacement or harmful drift. We prove a dynamic regret bound of $\tilde{O}(\log T)$ and validate the system on an edge-like single-GPU LoRA prototype running Qwen2.5-1.5B. Results demonstrate that WCA-UCB reduces model perplexity by $2.6$% to $8.6$% and uses up to $5.8\times$ fewer training-target switches than strong non-stationary baselines. These results highlight the necessity of convergence-aware scheduling for scalable local LLM personalization.
Chat is not available.
Successful Page Load