Gibbs Gradient Descent: A Langevin Approach to Optimization
Oussama Zekri ⋅ Anna Korba ⋅ Nicolas Boulle
Abstract
Many modern generative modeling and inference methods require optimizing objectives defined on parametric probability distributions. This leads to costly nested procedures, where each gradient step in first-order optimization techniques requires approximate sampling from the iterate parametric distribution. Gibbs gradient descent (GGD) was recently proposed as an efficient ``single-loop'' alternative by coupling sampling and optimization dynamics (Marion et al., 2025). However, its theoretical understanding remains limited, especially in the presence of particle approximations. This work provides a comprehensive analysis of GGD for the ideal infinite-particle system and its finite-particle approximation, explicitly quantifying the effect of particle noise via a sharp persistent $\mathcal{O}(1/n)$ variance term. We then identify a finite-observable structural class for the Gibbs gradient, which contains our motivating examples, and introduce a control-variate version of GGD that reduces the particle-induced variance at the observable level. Our results reveal a fundamental trade-off between optimization and sampling, driven by the stochastic error induced by finite particle approximations.
Chat is not available.
Successful Page Load