Correspondence Pruning by Iterative Structural Rectification
Guangwei Zhang ⋅ Gang Wang
Abstract
Correspondence pruning—separating sparse inliers from severe outliers in noisy feature matches—is a long-standing bottleneck for two-view geometry estimation in 3D vision. Existing learning-based pruners pursue this goal through ever more sophisticated context operators—local graphs, dynamic neighborhoods, or global attention—yet a controlled comparison reveals a striking phenomenon: when we freeze the structural assignment of these operators after a single forward pass, accuracy collapses by over $30\%$ mAP regardless of operator choice. This suggests that the limiting factor of current methods is not which operator is used, but the implicit assumption that structure can be committed in one shot. We argue that correspondence pruning should instead be cast as \emph{iterative structural rectification} (ISR), in which the geometric structure and the feature representation must \emph{co-evolve} layer by layer. Building on this view, we propose \textbf{ISR-Net}, a minimal architecture that pairs a layer-wise re-derived local-context operator (a dynamic $k$-NN graph in feature space) with a layer-wise re-derived global-context operator (soft assignment to learnable cluster prototypes), so that any gain over prior arts is attributable to the rectification \emph{schedule} rather than to operator novelty. Despite using deliberately off-the-shelf operators, ISR-Net attains state-of-the-art results on YFCC100M, SUN3D, and HPatches, surpassing the strongest attention-based competitor by $+8.2\%$ mAP@$5^\circ$ while running $35\%$ faster, indicating that \emph{how often structure is refined} matters more than the operator's complexity.
Chat is not available.
Successful Page Load