From Matching to Reasoning: Query-Aware Long Video Summarization
Abstract
Query-aware video summarization (QVS) aims to select a concise set of moments relevant to a textual query in the context of the entire video. Existing QVS methods are often driven by local query-segment relevance scoring, which is effective for retrieving query-matching moments but insufficient for composing coherent summaries over long videos. Long-form QVS is especially challenging because query-relevant evidence tends to be sparsely distributed across extended timelines, and satisfying a query may require combining complementary segments while avoiding redundancy under a strict summary budget. We propose QLVSumm, a framework that autoregressively generates a query-conditioned video summary, capturing inter-segment dependencies as well as the saliency of the segment and relevance to the query. We also introduce a large-scale benchmark for long-form QVS (LQVS) with open-vocabulary queries and budget-constrained extractive summaries, together with a complementary metric for evaluating query-aware summary disentanglement. Experiments show that QLVSumm achieves state-of-the-art performance and demonstrates stronger query-aware summary disentanglement.