Frank-LoRA: Federated Rank-Aware LoRA for Fine-Tuning Large Models
Saber Malekmohammadi ⋅ Kevin Kuo ⋅ Virginia Smith ⋅ Golnoosh Farnadi
Abstract
Fine-tuning large-scale foundation models across resource-constrained, distributed clients presents significant computational and communication challenges. While merging Low-Rank Adaptation (LoRA) with Federated Learning (FL) enables rank-based flexibility, the interdependence of the adaptation matrices ($A$ and $B$) often complicates optimization. In this work, we theoretically demonstrate that LoRA fine-tuning with a frozen A matrix acts as a stochastic approximation of full fine-tuning, where batch gradients are subject to a random, rank-dependent perturbation. Motivated by this insight, we propose freezing and periodically resampling $A$ matrices at the client level. This strategy decouples the $A$ and $B$ gradients, diminishes LoRA low-rank constraints and reduces both computational and uplink communication overheads by 50%, allowing for the reallocation of resources toward higher adaptation ranks. Furthermore, we show that in this frozen-$A$ regime, the adaptation rank $r$ directly governs the signal-to-noise ratio (SNR) of the perturbed gradients. Specifically, we find that the SNR increases linearly with $r$, implying that clients’ learning rates must be dynamically scaled with their adaptation ranks for higher utility. Building on these foundations, we introduce Frank-LoRA, a federated rank-aware fine-tuning strategy that utilizes frozen, resampled $A$ matrices and rank-dependent learning rates for clients. We also prove the convergence of the algorithm in FL settings with heterogeneous clients resources. Our experiments across benchmark datasets demonstrate that Frank-LoRA consistently outperforms state-of-the-art baselines with half the overhead on clients.
Chat is not available.
Successful Page Load