Contextual Flow Matching for High-quality Visual Content Generation
Abstract
Standard Flow Matching (FM) models typically employ a pointwise loss function, effectively assuming that the velocity vector at each spatial position is pointwise decomposed supervision given the conditioning. We argue that this Pointwise Independence Assumption is suboptimal for image generation, where strong local correlations exist across space. Neglecting these dependencies leads to artifacts such as structural inconsistency. In this paper, we introduce Contextual Flow Matching (CFM), a generalized framework that enforces consistency not just on individual pixels, but on the contextual relationships within the velocity field. We formulate the training objective using a set of generalized Context Operators---akin to convolution kernels---that abstract relationship constraints. Specifically, we instantiate a Differential Operator to capture high-frequency motion dynamics and an Aggregate Operator to stabilize low-frequency structural evolution. Extensive experiments on ImageNet demonstrate that CFM significantly accelerates convergence and enhances generation fidelity by explicitly modeling spatial dependencies.