Don't Pause! Every prediction matters in a streaming video
Abstract
Streaming video models should respond the moment an event unfolds, not after the moment has passed. Yet existing online VideoQA benchmarks remain largely retrospective. They pause the video at fixed timestamps, pose questions about current or past events, and score models only at those moments. This protocol leaves streaming outputs untested. To close this gap, we introduce SPOT-Bench, featuring multi-turn proactive queries that evaluate general streaming perception required by an always-on, real-time assistant. SPOT-Bench comes with Timeliness-F1, a consolidated metric that measures streaming outputs by their temporal precision and balanced coverage across the entire video. Our benchmark reveals that during streaming inference: (i) MLLMs detect events reliably in a stream but spam predictions unprompted; (ii) post-training MLLMs for silence reduces spamming but induces unresponsiveness; (iii) half of the streaming video expects no response, which we term dead-time - compute spent here does not affect response latency. These findings motivate Ctrl-SPOT, a training-free streaming controller for MLLMs, that retains their event perception while controlling their streaming outputs for balanced coverage. This establishes a strong baseline on SPOT-Bench.