Memory-Efficient Federated Fine-Tuning of LLMs via Block-wise Progressive Training
Qianyue Cao ⋅ Zongwei Zhu ⋅ Boyu Li ⋅ Yi Xiong ⋅ Zirui Lian ⋅ Xuehai Zhou
Abstract
Federated fine-tuning has become a dominant paradigm for privacy-preserving Large Language Model (LLM) adaptation. While integrating Parameter-Efficient Fine-Tuning (PEFT) reduces communication and computational costs, existing methods neglect that peak memory bottlenecks caused by full forward passes through the frozen LLM. This excludes low-memory devices, leading to data loss and suboptimal global performance. In this paper, we propose BP-FedPEFT, a framework utilizing progressive training to decompose the end-to-end computational graph, reducing peak memory to the block level. While enabling low-memory device participation, this paradigm incurs prolonged training latency and introduces growing memory burdens for deep-layer inputs, alongside suffering from cascading feature misalignment due to the absence of global supervision, leading to suboptimal model performance. To ensure efficiency, BP-FedPEFT employs functional-aware overlapping planning coupled with a local-global stability criterion to regulate training steps and communication rounds. To ensure effectiveness, we utilize depth-injected input synthesis and block overlaps to bridge the supervision gap, establishing valid optimization trajectories that align shallow representations with deep functional expectations. We establish theoretical convergence guarantees for BP-FedPEFT. Experiments on a heterogeneous testbed show that BP-FedPEFT supports diverse PEFT methods, reducing average memory usage by 44.7-76.2\%, accelerating training by 2.3-12.6$\times$, and improving accuracy by 3.2-5.8\% through inclusive participation.
Chat is not available.
Successful Page Load