Is Perception the Real Bottleneck? Measuring the Tool-Assisted Ceiling of Multimodal Reasoning
Abstract
Reliable visual reasoning requires not only the ability to draw correct inferences, but also access to the visual evidence needed to support them. We investigate how often failures of multimodal large language models arise from missing perceptual information rather than insufficient reasoning capacity. To quantify this distinction, we define a tool-assisted oracle ceiling: the best performance attainable by a frozen model when an optimal perception tool is selected independently for each example. Across five benchmarks, this oracle improves Qwen2-VL-7B by 18.4 percentage points, matching or exceeding gains obtained through fine-tuning on 522,000 samples. However, applying every available tool indiscriminately recovers only half of this improvement, indicating that effective routing—not tool availability—is the central challenge. Our sample-level analysis shows that perception tools correct 61.4% of baseline errors without training, compared with 50.9% corrected by scaling to Qwen3-VL-8B. Moreover, 77.1% of errors resolved through backbone scaling are also recoverable through better visual evidence alone. Motivated by these findings, we introduce PercepAgent, a training-free framework that routes queries by diagnosing missing evidence such as resolution, localization, depth, or region-level semantics. PercepAgent recovers 61% of the oracle headroom, generalizes across backbone generations without reconfiguration, and surpasses fine-tuned baselines on out-of-distribution benchmarks. These results highlight information access and faithful evidence integration as central bottlenecks in grounded multimodal reasoning.