An Educated Guess: Deriving Statistically Aligned Gradient Estimates for Zeroth-Order Optimization
Justus F Hübotter ⋅ Serge Thill ⋅ Marcel A. J. van Gerven ⋅ Nasir Ahmad
Abstract
As novel hardware architectures are explored for the purposes of machine learning, including neuromorphic chips, asynchronous systems, and ASICs, effective learning algorithms are highly desirable. While backpropagation is the standard for differentiation on standard hardware, its requirement for global synchrony and exact differentiability limits the exploration of exotic algorithms on non-traditional substrates and devices. Zeroth-order (ZO) optimization methods, such as SPSA and weight perturbation, offer a compelling alternative by estimating gradients through inference-only passes; however, these methods have historically suffered due to the "curse of dimensionality," where random perturbations become increasingly misaligned with the true gradient in high-dimensional spaces. In this work, we move beyond blind random noise by deriving zeroth-order search directions from activation geometry available during the forward pass. This leads to a family of Spectral Zeroth-Order methods, SZO-$\kappa$, whose directions solve local response-based gradient-guessing objectives under explicit stochastic assumptions. These methods spend perturbation budget on feature-supported directions that are more likely to produce informative loss responses. To further bridge the gap with first-order methods, we study covariance corrections and probe stabilizers, including search direction orthogonalization and bias corrections that reduce estimator error. Using multilayer and convolutional layers and networks, we analyze gradient alignment, bias-variance structure, and training stability of these methods relative to standard perturbation baselines and backpropagation.
Chat is not available.
Successful Page Load