A Benchmark for Omni-Modal Reasoning in Long Videos
Mohammed Irfan Kurpath ⋅ Jaseel M Kaithakkodan ⋅ Jinxing Zhou ⋅ Sahal Shaji Mullappilly ⋅ Mohammad Almansoori ⋅ Noor Ahsan ⋅ Beknur Kalmakhanbet ⋅ sambal shikhar ⋅ Rishabh Lalla ⋅ Jean Lahoud ⋅ Mariette Awad ⋅ Fahad Shahbaz Khan ⋅ Salman Khan ⋅ Rao Anwer ⋅ Hisham Cholakkal
Abstract
Long-form omni-modal video understanding requires models to integrate vision, speech, and ambient audio with coherent long-context reasoning. Existing video benchmarks often trade off temporal scale, modality coverage, open-ended interaction, and interpretable scoring. To address this gap, we introduce $\textbf{LongShOTBench}$, a long video evaluation benchmark designed around three coupled goals: holistic omni-modal integration, intent-driven open-ended interaction, and rubric-level diagnosis. It constructs single- and multi-turn questions from practical viewing scenarios, involving systematical tasks for probing visual, speech, ambient-audio, temporal, and cross-modal reasoning. Each item includes a reference answer and a weighted criterion-level rubric, enabling evaluation to identify which perceptual facts, temporal links, modality-grounding requirements, and reasoning steps are satisfied or missed. All samples are manually verified and corrected to improve grounding, clarity, and rubric reliability. We also introduce $\textbf{LongShOTAgent}$, a training-free omni-modal evidence-seeking agent that couples full-video preprocessing with targeted retrieval, query-adaptive segment refinement, and explicit claim verification over visual, speech, and non-speech audio evidence. Its iterative search-refine-verify loop exposes intermediate evidence and lets modality-specific specialists re-analyze relevant moments before answering. We perform comprehensive evaluation of 105 video-capable models spanning open-source omni-modal models, vision-language systems, audio LLMs, agentic pipelines and closed-source APIs. Across this broad evaluation, current MLLMs remain far from saturating LongShOTBench, while our LongShOTAgent emerges as the strongest training-free system, reaching 66.64\% overall. By releasing the benchmark, leaderboard, and agentic method, our work provides the community with a shared, interpretable testbed for evaluating and advancing long-form omni-modal video reasoning.
Chat is not available.
Successful Page Load