How I learned to stop worrying and love StopGrads: Stationarity, Convergence, and a case study on Flow Map Learning
Mark Goldstein ⋅ Max Shen ⋅ Zichu Wang ⋅ Aahlad Manas Puli ⋅ Rajesh Ranganath
Abstract
Stopgrads are widely used in training machine learning models, but stopgrads can alter the gradient, stationary points and convergence guarantees of the original objective, which can make stopgrad training theoretically ungrounded. We introduce a \textit{stopgrad regression principle}, which identifies a general template for stopgrad objectives with a closed-form characterization of stationary points, and unifies stopgrad objectives for flow maps, reinforcement learning, and diffusion samplers. We provide theoretical grounding for optimizing stopgrad flow map objectives by characterizing stationary points and proving convergence results. Remarkably, we show that the learned flow map has a closed-form expression composing the initial flow map and the true flow map, under functional semi-gradient flow. We use our stopgrad regression principle to propose novel modified stopgrad placements for flow map objectives which speed up training by ${\sim}2.5\times$.
Chat is not available.
Successful Page Load