Vision Correlators: Correlation-Driven Visual Understanding with Hypergraphs
Abstract
Modern visual backbones, despite their architectural diversity, largely reduce relational reasoning to pairwise message passing on graphs and therefore primarily capture 2-order interactions. However, visual understanding often depends on latent yet crucial higher-order semantic correlations among multiple regions, which cannot be adequately expressed by pairwise modeling alone. We first formalize this limitation by proving an irreducibility result: on a fixed vertex set, there exist nonlinear hypergraph message passing layers that cannot be represented exactly by any graph message passing layer unless the hypergraph satisfies highly restrictive structural conditions. This shows that higher-order correlation modeling offers a fundamental expressive advantage rather than a simple approximation to graph-based backbones. Therefore, we propose Vision Correlator (ViC), a general-purpose visual backbone built on multi-order hypergraphs. ViC consists of correlation induction and correlation propagation. In the correlation induction stage, we propose an Anchor-based Hypergraph Generation strategy that uses anchors as semantic centers and incorporates topological information to generate scalable hypergraph structures. In the correlation propagation stage, we further design Differential Mixed Aggregation to achieve efficient hyperedge representation and vertex feature updating, and introduce hyperedge dropout to improve the stability of structure learning. Extensive experiments with isotropic and pyramid variants show that our ViC achieves a better accuracy-efficiency trade-off than strong Transformer and graph-based baselines. Specifically, compared to the Vision Transformer baseline, ViC achieves up to 76\% parameter reduction, 93\% FLOPs reduction, and delivers up to a 5.1\% increase in accuracy on ImageNet.