ViDiC: Video Difference Captioning
Abstract
The rapid advancement of controllable video editing and agent-driven video generation has exposed a critical bottleneck: the inability of existing Multimodal Large Language Models to accurately perceive and articulate fine-grained differences between videos. Whether diagnosing editing fidelity or detecting training data quality issues, explicit comparative perception serves as an essential infrastructure. To address this capability gap, we introduce the ViDiC task and its corresponding ViDiC-1K benchmark. ViDiC-1K comprises 1,000 curated video pairs annotated with 3,720 fine-grained comparative checklist items, rigorously structured across seven dimensions: subject, style, background, camera work, motion, position, and playback techniques. To ensure reliable evaluation and prevent model hallucination, we propose a novel dual-checklist framework that assesses the accuracy of both similarities and differences separately based on an LLM-as-a-Judge protocol. Extensive experiments on 16 representative MLLMs reveal significant limitations in their comparative reasoning and dynamic difference perception. We hope ViDiC-1K serves as a foundational benchmark to advance continuous multi-video captioning and edit awareness in multimodal intelligence.