VideoSleuth: A Narrative-Centric Agentic System with Video-Audio Native MLLM for Long-Form Video Understanding
Conglin Li ⋅ Yang Li ⋅ Qi Zhang ⋅ Guanhua Chen ⋅ Yun Chen ⋅ Weifeng Ge
Abstract
Long-form video understanding (LVU) requires models to reason over hour-long narratives while preserving fine-grained spatiotemporal details across extended temporal horizons. Although recent multimodal large language models (MLLMs) have achieved strong performance on short-video tasks, their reasoning capabilities degrade substantially as video duration and information density increase. We present VideoSleuth, a narrative-centric agentic framework for LVU built on video- and audio-native MLLMs. VideoSleuth organizes its memory around the four narrative elements-time, place, person, and event-and equips the agent with tools explicitly designed to construct and query these elements within a ReAct-style reasoning loop. Unlike prior LVU agents that rely on dense offline preprocessing or image-only VLMs, VideoSleuth performs perception on demand using a video-audio native perceptual backbone, jointly leveraging visual and auditory cues to support dialogue-based identity grounding, long-range event localization, and fine-grained evidence retrieval. On four standard LVU benchmarks, VideoSleuth consistently outperforms prior on-demand grounding agents and achieves competitive accuracy with the strongest dense-preprocessing baseline, while consuming at least $5.2\times$ fewer end-to-end tokens than DVD-style pipelines spend on offline preprocessing alone.
Chat is not available.
Successful Page Load