FuRA: Full-Rank Parameter-Efficient Fine-Tuning with Spectral Preconditioning
Yequan Zhao ⋅ Ruijie (Ray) Zhang ⋅ Liyan Tan ⋅ Niall Moran ⋅ Tong Qin ⋅ Zheng Zhang
Abstract
Both full fine-tuning (Full FT) and parameter-efficient methods like LoRA add weight updates without regard to the spectral structure that pretraining has established. This allows noisy gradients from a small fine-tuning distribution to freely perturb the robust features learned through pretraining. We first identify *spectral preconditioning* as the key missing ingredient: reparameterizing each weight $\mathbf{W}$ through its full-rank SVD and freezing one singular basis confines every update to the pretrained column space, yielding a preconditioned optimizer that outperforms unconstrained Full FT at the same parameter count. To make this insight practical, we propose FuRA (**Fu**ll-**R**ank **A**daptation), which factorizes $\mathbf{W}$ via a block tensor-train decomposition $\mathbf{W}=\mathbf{L}\mathbf{S}\mathbf{R}$: the large core $\mathbf{L}$ is frozen at the pretrained block-wise SVD basis while only the small core $\mathbf{R}$ and per-block singular values $\mathbf{S}$ are trained. This single design choice simultaneously delivers full-rank spectral preconditioning, full-rank update capacity, and parameter, step time, memory efficiency on par with LoRA. FuRA outperforms Full FT on LLM fine-tuning ($+1.37$ on LLaMA-3-8B commonsense reasoning), LLM math reinforcement learning, and VLM visual instruction tuning. The 4-bit quantized version QFuRA also outperforms QLoRA.
Chat is not available.
Successful Page Load