Dense Flow from Adaptive Correspondence
Abstract
Dense optical flow estimates correspondence from every pixel to the next frame. It must be precise and robust, which couples two choices: what units to match, and what evidence to match them. Recent work has improved the evidence with stronger features, cost volumes, attention, refinement, and generative priors. The unit, however, usually remains fixed: motion is still estimated at dense grid points. We propose Adaptive Correspondence Tokens (ACT), which adapts correspondence units through a differentiable pixel-to-token assignment. ACT learns soft maps that group pixels into tokens, build descriptors, match and refine motion in token space, and project token motion back to pixels for dense flow. Trained end to end with the flow loss, grouping and correspondence are discovered jointly: grouping makes matching easier, while matching shapes grouping. Support can expand over coherent regions and contract near boundaries or fine detail. In zero-shot cross-dataset evaluation, ACT achieves the best KITTI EPE, second-best KITTI Fl, and second-best Sintel Final EPE among published methods, while remaining competitive under benchmark fine-tuning. These results show that adaptive units make dense flow easier: matching where the image gives reliable evidence improves transfer while preserving dense, boundary-aware prediction.