Fixed-Point Rates for Self-Distillation of Flow Maps
Abstract
Flow and diffusion models generate high-quality samples, but inference typically requires many sequential neural-network evaluations. Flow-map models address this bottleneck by learning finite-time transitions directly, enabling generation in one or a few evaluations. Previous work introduced Lagrangian, Eulerian, and progressive self-distillation rules with training targets derived from the model's own predictions. Does reconstructing a candidate flow map from its own predictions correct errors or reinforce them? With the learned velocity field held fixed, we view each rule as an operator on candidate flow maps and study its iterates. By composing two half-length transitions, progressive reconstruction converges globally and geometrically in a weighted norm. This requires spatial Lipschitz bounds and initial errors relative to the teacher flow that vanish uniformly at least quadratically as transitions become shorter. Every reconstruction reduces the worst-case weighted error by at least half, although it can amplify error in the ordinary supremum norm. Lagrangian reconstruction has a factorial convergence guarantee, whereas Eulerian reconstruction admits no general same-order Lipschitz stability bound. Under the stated assumptions, Lagrangian and progressive reconstruction yield two-sided bounds on the error to the frozen teacher flow in terms of the full weighted residual norm. For progressive reconstruction, analogous mean-square bounds hold over teacher-generated states, with a supremum over time pairs. A two-dimensional experiment illustrates how reconstruction reduces errors over short intervals and how its changes indicate the remaining error. Lagrangian reconstruction gives faster initial improvement in this example.