Reasoning-based Spatial Prior (RSP): Learning Spatial Priors from Multimodal LLMs for Object Detection
Abstract
Grounding natural language queries to target objects in images requires both high-level reasoning about intentions and precise spatial localization. Multimodal large language models (MLLMs) excel at reasoning about which object is desired given an implicit intention query but turning this understanding into precise bounding box predictions remains challenging. Conversely, specialized detectors such as GroundingDINO offer robust localization capabilities but lack the inherent capacity to understand implicit intentions. Existing hybrid methods couple these components only at the semantic level, such as by passing object names or a single global token to the detector, thereby discarding the rich, dense spatial signals encoded in MLLM attention maps. We propose Reasoning-based Spatial Prior (RSP), a reasoning-based detection framework that converts the MLLM’s internal attention into a learnable spatial prior and injects it directly into the detector. Concretely, we learn input-adaptive gating over attention heads, supervise the fused spatial prior with a SoftIoU localization loss, and integrate it into GroundingDINO via a prior-guided deformable fusion module that biases sampling toward regions highlighted by the reasoning process. Our method consistently outperforms strong MLLM-only, detector-only, and hybrid baselines on the EgoIntention and RIO benchmarks, achieving state-of-the-art performance with substantial gains across both frequent and uncommon categories.