SaVE: Surface-Aware Visual Evidence Selection for Zero-Shot 3D Question Answering
Abstract
Zero-shot 3D question answering enables pretrained vision-language models (VLMs) to reason about scenes from multi-view images without task-specific 3D training. Its central challenge is to fit spatially distributed answer evidence into a limited visual context. Neighboring views often repeat the same physical surfaces, so redundant observations can displace complementary evidence needed to answer a question. We present SaVE, a surface-aware visual evidence selection method that addresses this coverage-redundancy trade-off. SaVE uses cross-view surface correspondence to evaluate how much each observation adds to the evidence already retained. It prioritizes relevant, reliable observations of surfaces that are missing or poorly represented, while reducing repeated evidence from well-covered surfaces. This produces a compact visual context for a frozen VLM. Experiments on ScanQA, SQA3D, and VSI-Bench show answer-quality gains across a range of models and context budgets. SaVE also retains broader physical surface evidence and approaches full-context ScanQA performance with lower peak memory.