LookThere! Sparse Vision by Reinforced Selection
Abstract
Essential visual information can be sparse, yet most visual computation remains dense. Though only a fraction matters, vision transformers process images as uniform sets of tokens, scaling cost with their number. As these models grow stronger, they grow larger, and the cost of computing everything climbs higher. Adaptive computation has shown that we can get by with less, but existing methods struggle at extreme sparsity and are difficult to specialize to a task. We address this limitation by introducing LookThere, training an efficient selector of locations and an expressive extractor of representations to derive meaning from images without ever processing the full high-resolution input. We do so by learning the selector and extractor end-to-end with actor-critic reinforcement learning. The selector learns where to look, the extractor what to see, together saving computation by selecting only what is worth processing for a given task. We show that LookThere selects the task-specific input, excelling at sparse recognition in high-resolution settings (traffic signs, billiards), and maintaining accuracy with as little as 0.2% of the input. It also generalizes across tasks and models, including global recognition (ImageNet classification), local recognition (ADE20K segmentation), zero-shot classification (by distillation), and regression (class agnostic enumeration). Across all settings, LookThere surpasses state-of-the-art selection, providing a general and scalable framework for specialized and efficient adaptive computation.