On Locality and Length Generalization in Visual Reasoning
Abstract
A striking feature of human vision is that it acquires information through local, high-resolution glimpses, whereas most vision models encode an image globally in one pass. A natural question therefore is whether local, sequential vision models may provide any foundational computational benefits in addition to being biologically more plausible than global models. We investigate this question from the perspective of visual state tracking and length generalization. We study the behavior of vision-language models trained on simple vision tasks where a model is required to aggregate local information distributed across an image and update the state of a system. Length generalization asks whether an algorithm trained on shorter sequences can extrapolate to longer sequences at test time. Our experiments reveal that, similar to language models, vision models can learn to exploit global shortcuts and thereby fail to generalize over task length and spatial complexity. We also show that recurrent vision policies based on strictly local perception can mitigate these failures, thereby allowing models to generalize on these tasks. Our results show that local attention may be an essential overlooked requirement for robust compositional generalization.