Better Source, Better Flow: Learning Condition-Dependent Source Distribution for Flow Matching
Junwan Kim ⋅ Jiho Park ⋅ Seonghu Jeon ⋅ Seungryong Kim
Abstract
Flow matching has recently emerged as a promising alternative to diffusion-based generative models, particularly for text-to-image generation. Although flow matching places no restriction on the source distribution, most existing systems still inherit a standard Gaussian from diffusion models, and the source is rarely treated as an optimization target at this scale. Recent works have begun to revisit this choice through condition-dependent or learned sources, yet evidence that such designs are effective in modern text-to-image systems—with high-dimensional latents and tightly integrated conditioning—remains limited. In this work, we study condition-dependent source distributions for flow matching along three axes: _why_ source learning helps, through the lens of the intrinsic variance term in the flow-matching objective; _how_ to make it work in modern text-to-image systems, where variance-only regularization and directional source—target alignment are critical for stable end-to-end training; and _when_ it is most beneficial, by connecting source design to recent representational advances in generative modeling and identifying target representation regimes in which learning the source yields the largest gains. Extensive experiments across multiple text-to-image benchmarks, backbones, and scales demonstrate that principled source design yields consistent and robust improvements—including up to $\mathbf{3.01\times}$ faster convergence in FID and $\mathbf{2.48\times}$ in CLIP score—and outperforms representative prior conditioned-source and condition-aware coupling methods.
Chat is not available.
Successful Page Load