PixSearch: Region-Grounded Retrieval for Knowledgeable Large Multimodal Models
Abstract
Visual Question Answering (VQA) often requires external knowledge about the specific entity or region relevant to a question, not just recognition from the full image. This is especially critical for egocentric images from wearable devices, where small, off-center, or long-tail objects make global retrieval noisy, yet region-level knowledge is needed for factual reasoning and situated decision making. Existing Multimodal Retrieval-Augmented Generation (MM-RAG) systems rely on full-image retrieval, text-only queries, or modular detector/segmenter pipelines, introducing irrelevant evidence, cross-module errors, and latency. We propose PixSearch, an end-to-end segmenting Large Multimodal Model (LMM) that unifies region-level visual grounding with retrieval-augmented reasoning. During autoregressive generation, PixSearch learns when to retrieve, how to route queries across text, whole-image, and region-level retrieval, and how to crop the target entity as a visual query using a jointly trained mask decoder. A two-stage supervised fine-tuning regimen with search-interleaved supervision teaches retrieval triggering, query selection, and evidence-grounded answer generation while preserving segmentation ability. On egocentric and entity-centric VQA benchmarks, PixSearch achieves a 19.7% relative accuracy gain on CRAG-MM over whole-image retrieval, with particularly strong gains on egocentric images, while remaining competitive on broader VQA and text-only QA tasks. It is also 44% faster than pipeline MM-RAG frameworks in terms of inference latency.