CineMME: Benchmarking Fine-Grained Perception and Plot Reasoning in Multimodal Large Language Models
Jin Liu ⋅ Dawei Du ⋅ Yexiang Liu ⋅ Sijie Zhu ⋅ Fan Chen ⋅ Junxian Duan ⋅ Huaibo Huang ⋅ Ran He ⋅ Longyin Wen ⋅ Zhenfang Chen
Abstract
Cinematic understanding requires more than recognizing visible events: models must acquire dialogue evidence, bind it to speakers, and use it to infer narrative meaning. Existing video benchmarks evaluate pieces of this problem, including video question answering, audio-visual reasoning, and temporal grounding, but rarely test whether these abilities form a coherent cinematic evidence chain within the same clips. As a result, a model may answer a movie question correctly while failing to localize the relevant utterance, identify the speaker, or integrate speech with visual context. We introduce **CineMME**, a human-verified benchmark designed to diagnose speech-grounded cinematic reasoning, supported by approximately $1,600$ expert person-hours of annotation. CineMME contains $975$ curated movie clips and two linked tracks: *CineMME-Grounding*, with $13,554$ timestamped speech segments and $33,083$ actor face boxes for temporal dialogue localization, transcription, and active-speaker grounding; and *CineMME-Reasoning*, with $1,600$ multiple-choice and $316$ open-ended questions spanning perception, audio, narrative, social, and cinematic understanding. We evaluate $20$ MLLMs on reasoning and $10$ omni-modal models on grounding. The results reveal a pronounced who-what-when decoupling: Gemini-3 Pro localizes dialogue in time 68.33% tIoU) and partially transcribes it (WER 29.00%), yet reaches only 13.24% sIoU for active-speaker grounding while achieving 65.88% reasoning accuracy. Modality ablations further show that vision is not uniformly beneficial: it improves perception-heavy questions but can hurt dialogue-centric reasoning. CineMME provides a compact diagnostic testbed for measuring whether multimodal video models truly connect dialogue, speakers, visual evidence, and narrative meaning, rather than succeeding through disconnected cues.
Chat is not available.
Successful Page Load