PIVOT: A Unified Agentic Framework for Streaming Long-Video Understanding
Abstract
Streaming long-video understanding is a realistic setting for always-on video assistants, where visual input arrives continuously and users may ask questions at arbitrary times. We study streaming long-video question understanding (QA) under a unified temporal formulation in which multiple questions are revealed over time and may concern the past, present, or future relative to their query times. The model receives no task-type labels and must operate under strict sequential access and bounded shared memory, deciding whether each active question is already answerable or should remain open for future evidence. This differs from prior streaming settings that often focus on short streams, prefix-answerable queries, separated temporal categories, or response timing without long-horizon shared memory. To evaluate this formulation, we propose Ref2Stream (Reference-to-Stream), a general benchmark-conversion protocol that turns temporally grounded offline video QA benchmarks into streaming QA benchmarks. Instantiated on LVBench, Ref2Stream yields LVBench-Ref2Stream, which mixes past, present, and future questions under strict sequential access and evaluates both answer accuracy and answer latency. Existing streaming video methods typically improve efficiency or response timing through predefined compression, retrieval, or response policies, while offline video agents perform adaptive evidence gathering but assume access to the full video and target question. We introduce PIVOT (Planning over Incremental Video Observations for Timely Answering), a unified agentic framework that incrementally updates bounded memory from the observed stream and uses a planner-controlled action loop to retrieve relevant memory, analyze current evidence, reflect on candidate answers, and decide whether to answer or defer. Our work suggests that agentic reasoning is a promising direction for building streaming long-video systems that can adaptively gather evidence and answer only when sufficient information has been observed.