Towards Identifying Dominant Low-Rank Subspaces in Zeroth-Order Fine-Tuning
Abstract
Adapting Large Language Models (LLMs) to downstream tasks is increasingly bottlenecked by the staggering memory overhead of first-order (FO) backpropagation. Zeroth-order (ZO) optimization has emerged as a compelling, memory-efficient alternative; however, it suffers from the curse of dimensionality, which introduces prohibitively high variance in gradient estimation. While LLM gradients empirically reside in low-rank subspaces, existing state-of-the-art ZO methods typically rely on randomly generated low-rank subspaces, which often fail to align with the underlying dominant gradient manifold, fundamentally limiting optimization efficiency. To bridge this gap, we propose AHZO, an efficient low-rank ZO fine-tuning algorithm that dynamically identifies dominant subspaces via the Average Historical Gradient (AHG). Motivated by the directional coherence of optimization trajectories, AHZO leverages the average historical ZO gradient over a period as a principled proxy for the true gradient. By applying singular value decomposition to AHG matrices, AHZO distills principal spectral signatures to construct low-rank bases that tightly align with the true dominant subspaces. Theoretically, we prove that AHZO significantly mitigates estimation variance, yielding superior convergence guarantees compared to random-subspace ZO methods. Extensive evaluations across diverse LLM architectures and benchmarks demonstrate that AHZO consistently outperforms existing ZO baselines, achieving performance highly competitive with memory-intensive FO fine-tuning.