The Cost of Mismatch: Noise Amplification in Zeroth-Order Reinforcement Learning
Lianmin Chen ⋅ Junbin Qiu ⋅ Chenxing Wei ⋅ Yao SHU ⋅ Kun He
Abstract
Zeroth-order optimization (ZOO) has emerged as a promising approach for fine-tuning Large Language Models (LLMs) when gradient access is unavailable or memory-prohibitive. However, ZOO relies on estimating gradients via finite differences between two function evaluations at perturbation scale $\mu$, making it inherently noisier than its first-order counterpart. This raises a fundamental question: \textit{under what conditions does ZOO yield reliable gradient estimates?} In this work, we provide a rigorous variance-theoretic answer through a generic noise-amplification framework and demonstrate that the reliability of ZOO gradient estimations is structurally determined by whether the two finite-difference evaluations are performed on \textit{matched data}. Mismatched evaluations, as induced by on-policy Reinforcement Learning (RL), produce a $1/\mu^2$ oracle-noise amplification that sets an irreducible noise floor, whereas matched evaluations, as realized by offline optimization methods such as Boltzmann Targeted SFT (BOLT), eliminate this amplification entirely via noise cancellation. We further derive convergence guarantees with explicit constants that cleanly separate the two regimes. Controlled experiments on math reasoning tasks validate these findings, using ZooBOLT and ZooGRPO --- zeroth-order adaptations of BOLT and GRPO --- as representatives of the matched- and mismatched-data regimes respectively. On GSM8K with Qwen2.5-1.5B, ZooBOLT recovers $96.1\%$ of first-order BOLT's performance gains, while ZooGRPO exhibits near-zero learning despite its first-order counterpart achieving strong results. This qualitative separation persists across harder benchmarks (MATH), larger model scales (Qwen2.5-7B), and different model variants (DeepSeek-R1-Distill-Qwen-1.5B). Additionally, ZOO methods reduce peak memory to near-inference levels, validating its practicality as a memory-efficient fine-tuning paradigm.
Chat is not available.
Successful Page Load