LocalAgent: Collaborative Agentic Verification for Fine-Grained Instance-Level Consistency
Abstract
Multimodal large language models (MLLMs) have made rapid progress in general visual understanding, yet they still struggle with fine-grained instance-level consistency verification, where the goal is to determine whether two images depict the exact same physical object under varying viewpoints, illumination, and backgrounds. This task requires both a global understanding of the target instance and a precise examination of subtle local evidence, since semantically similar objects may differ only in small discriminative regions. To address this challenge, we propose \textbf{LocalAgent}, a collaborative agentic framework that integrates the complementary strengths of large and small models. Specifically, the MLLM first reasons from a global instance-level perspective to identify suspicious local regions that may determine the final verification result. These accurately localized regions are then delegated to a lightweight Local Evidence Verifier (LEV) toolkit for fine-grained appearance and geometric matching. The LEV toolkit consists of \textbf{Local Appearance Consistency (LAC)}, which provides illumination- and color-robust similarity estimation, and \textbf{Local Geometric Correspondence (LGC)}, which establishes reliable spatial correspondence under viewpoint changes. Rather than directly replacing the MLLM's reasoning, LEV returns structured similarity scores and comparative evidence in a form that the MLLM can effectively interpret for final decision-making. This design enables a coordinated large-small model collaboration: the MLLM contributes global semantic reasoning and region-level suspicion localization, while specialized local verifiers provide precise, invariant, and interpretable evidence for ambiguous cases. We further introduce an answer-dominant reinforcement fine-tuning (RFT) strategy that enables the policy to learn when verification is necessary. Experiments show that LocalAgent achieves a \textbf{+4.8\%} average improvement over direct-inference baselines, effectively narrowing the gap with closed-source models, while preserving the general multimodal capabilities of the underlying MLLM.