Improving Quantized Zeroth-Order Optimization through Reconstructed Low-Rank Structures
Abstract
Fine-tuning Large Language Models under stringent memory constraints remains a significant challenge. Recently, Zeroth-Order (ZO) optimizers have been employed to fine-tune quantized models, avoiding backpropagation and further reducing memory usage. However, quantized ZO optimizers often underperform compared to their full-precision counterparts, especially at low bit-widths. We find that existing quantized ZO methods fail to exploit the inherent low-rank structure of gradient updates, thereby limiting their effectiveness. To address this, we propose the Low-rank Quantized Zeroth-Order optimizer (LQZO), which enhances quantized ZO fine-tuning by leveraging the low-rank structure inherent in gradients. Our method introduces three key innovations: (1) A low-rank parameterization of gradient perturbations using structured matrix factorization. (2) A Binary-Aggregated Gaussian perturbation strategy that employs a Rademacher distribution to constrain higher-order moments and tighten the variance bound of gradient estimates, with no additional computational overhead. (3) An Isotropy-Enhanced Estimation technique that reshapes quantization scale groups into more isotropic structures to diversify exploration directions. We provide a theoretical convergence guarantee for LQZO. Extensive experiments demonstrate that LQZO consistently outperforms quantized ZO optimizers, achieving average performance gains of 2.0%--3.3% across quantization precisions. While extreme quantization inevitably leads to information loss, LQZO significantly improves the trainability of INT2 models and narrows the performance gap compared to prior quantized ZO methods. Furthermore, we demonstrate the cross-modality versatility of LQZO by validating its superior performance on Vision Transformers.