Geometry-Aware Zeroth-Order Optimization for Fine-Tuning Quantized LLMs
Shaocong Ma ⋅ Weidong Cai ⋅ Heng Huang
Abstract
Zeroth-order optimization (ZOO) is a promising approach for memory-efficient fine-tuning of Large Language Models (LLMs). However, scaling ZOO to larger models on a single GPU necessitates aggressive quantization, which often destabilizes training. A primary cause of this instability is that existing methods treat strictly positive quantization scale parameters as Euclidean variables, potentially leading to boundary violations and optimization difficulties. To address this issue, we propose Hyper-Octant Zeroth-order Optimization (HoZO), a geometry-aware framework that formulates fine-tuning as Riemannian optimization on the positive orthant manifold. By constructing a novel geodesically complete metric and deriving the closed-form exponential map, HoZO inherently enforces positivity constraints without bias-inducing projections. On the theoretical side, we prove that HoZO achieves the optimal oracle complexity for zeroth-order methods. On the empirical side, HoZO outperforms baselines across six downstream tasks and three model sizes. Notably, it achieves a $34\times$ memory reduction compared to first-order fine-tuning, enabling the stable fine-tuning of a 4-bit quantized Llama-2-70b model on a single 48GB GPU.
Chat is not available.
Successful Page Load