Beyond Spatial and Temporal Priors: A Generalizable Approach for Dense Correspondence Matching
Abstract
Dense correspondence matching has historically been bounded by simplifying spatio-temporal priors, such as smooth motion and rigid geometry. While effective for classical tasks, these foundations collapse at the frontier of Vision-Language-guided Image Editing and Generation (VL-IEG), where transformations yield correspondences that are perceptually obvious yet physically discontinuous. To transcend spatio-temporal priors, we draw inspiration from human cognition: learning a generalized model of identity-preserving visual consistency rather than relying on physical constraints. To achieve this, we introduce FreeMatching, a unified framework that extracts identity-preserving knowledge from foundation models and diverse datasets. The model's capability is forged through a two-stage training paradigm: supervised pre-training on diverse annotated datasets, followed by weakly-supervised refinement on data without dense annotations. Experimentally, FreeMatching not only achieves performance competitive with SOTA methods on classical benchmarks but also establishes the first strong baseline for the challenging VL-IEG task. Furthermore, we demonstrate its utility as a quantitative metric for evaluating identity-preserving consistency in generative editing, aligning closely with human judgment.