Beyond Flat Frames: Hierarchical Graph Reasoning for Long Video Understanding
Abstract
Recent studies have shown that leveraging the reasoning and analytical capabilities of large language models (LLMs) for long video understanding has become a promising approach. However, these methods are constrained by a structural limitation: representing video as a flat frame sequence makes it difficult to model the temporal structure and logical dependencies of events, which in turn hampers long-range reasoning and risks overlooking important information. To address this limitation, we propose LVGraph, a framework that models video content using a coarse-to-fine hierarchical graph. LVGraph constructs a semantic hierarchy where high-level graphs characterize the global event structures within the video, while low-level graphs capture fine-grained, query-relevant information. Instead of linear scanning, reasoning is performed via a query-aware traversal of this graph, adaptively identifying and retrieving the most salient keyframes by navigating the nodes and edges most relevant to the query. Finally, we leverage a Vision Language Model (VLM) to augment the keyframe information, and subsequently answer the question by reasoning over the graph and the enriched keyframe captions. Extensive experiments on long video understanding benchmarks confirm the effectiveness of our method. LVGraph significantly outperforms existing LLM-based state-of-the-art methods on EgoSchema, NExT-QA, and long-duration segments (averaging 44 minutes) of the Video-MME benchmarks.