Seeing with Intent: Target-Guided Visual Foveation for Multimodal Reasoning
Abstract
Multimodal large language models (MLLMs) have made rapid progress in visual understanding and reasoning. Yet visual reasoning is not simply a perception-then-reasoning process: observation and inference proceed in alternation, and the visual information needed at each stage depends on the model's current reasoning intent. Many existing active-vision methods convey primarily spatial decisions to perception. Operations such as cropping or zooming specify where and at what scale to look, but do not allow the visual intent formed at the current reasoning step to directly shape how visual information is represented. We introduce Target-Guided Visual Foveation (TGVF), which feeds the model's current visual intent back into visual processing as reasoning unfolds. The model specifies this intent as a natural-language visual target, and TGVF uses the target's contextual hidden states to condition the visual representation through a bidirectional target--vision adapter. The resulting target-conditioned visual evidence is returned to the same reasoning trajectory to support subsequent inference. To train this interface, we construct TGVF-50K, a 47,921-trajectory training corpus, of which 40,535 contain visual targets and image-grounded evidence descriptions, and train the adapter to provide the requested visual information in a form usable by the base model. We then use reinforcement learning to integrate this capability into the model's reasoning process without relying on cold-start SFT. We instantiate TGVF on Qwen3-VL-8B-Instruct and evaluate it on six visual-reasoning benchmarks. Controlled analyses validate that TGVF provides visual evidence that is usable by the base model and responsive to the expressed intent. TGVF raises the six-benchmark average from 52.19% to 57.72%, while combining it with Crop further improves the average to 61.99%, showing that target-conditioned visual processing complements spatial localization.