From Finding to Linking: Benchmarking and Advancing Cross-Long-Video Reasoning for Multimodal LLMs
Abstract
Reasoning across multiple long-form videos requires models to find sparse evidence within each video and link related entities, events, and narratives across streams. Existing long-video benchmarks mainly evaluate single-video understanding, while multi-video benchmarks typically use short clips. We introduce CLoVR-Bench, a comprehensive benchmark for Cross-Long-Video Reasoning with 2,000 expert-annotated QA pairs over 400 long-form videos. Its three-level taxonomy covers Comparative Analysis, Tracking and Retrieval, and Integrated Reasoning, spanning 14 tasks and 36 subtasks that diagnose both intra-video localization and inter-video alignment. Our evaluation of 13 representative MLLMs shows substantial gaps across this hierarchy, with performance degrading from localized comparison to long-range retrieval and integrated cross-video reasoning. We further propose HOLMES, a training-free framework that formulates cross-long-video reasoning as evidence-slot filling over a typed binding graph. HOLMES plans option-discriminating evidence slots, localizes evidence with risk-conditioned policies, records coverage certificates for failed searches, verifies cross-video bindings visually, audits constraints, and returns answers through hard graph-readout gates. Experiments show that HOLMES improves both open-source and closed-source backbones, outperforms adapted long-video baselines, and yields the largest gains on tasks requiring long-range evidence search and cross-video binding. Together, CLoVR-Bench and HOLMES provide a diagnostic testbed and an interpretable baseline for future cross-long-video reasoning.