Cross Flow: One-Step Generation Across Latent and Pixel Spaces
Xiyuan Wang ⋅ Xiao Zhang ⋅ Yang Li ⋅ Ruoxi Jiang ⋅ Zhao Zhong ⋅ Liefeng Bo ⋅ Muhan Zhang
Abstract
Diffusion and flow-matching models are typically constrained to a single manifold, where the noise prior, intermediate trajectory, and prediction target share the same dimensionality. While Latent Diffusion Models (LDMs) mitigate computational costs by operating in compressed spaces, they still rely on a decoupled, fixed decoder to map latents back to pixels. We introduce CrossFlow, a generative paradigm that unifies these stages by allowing the noise prior and final output to reside in different spaces. By deriving a novel cross-space objective, our framework enables a single model to map directly from a noisy latent to a high-resolution image in a single function evaluation (NFE). CrossFlow serves a dual purpose: it acts as a high-fidelity one-step generator and functions as an enhanced decoder for existing LDM pipelines, capable of refining imperfect latent estimations during the reconstruction process. On ImageNet-1k ($256 \times 256$), CrossFlow achieves a state-of-the-art 1.62 FID with only one NFE. Our results demonstrate that cross-space flow objectives provide a scalable and theoretically grounded framework for unifying latent generation and pixel-space decoding.
Chat is not available.
Successful Page Load