DriveStreamBench: Evaluating User-Conditioned Watch-and-Notify in Streaming Driving Video
Abstract
Vision-language models (VLMs) have made steady progress as in-vehicle driving assistants, yet existing evaluations largely remain offline, focusing on perception and short-term reasoning over pre-recorded driving scenes. A complementary capability is still underexplored: whether a model can continuously monitor a causal driving video stream and issue a calibrated alert when a user-specified condition is met. Existing driving benchmarks mainly assess reactive understanding, while recent streaming video benchmarks are developed largely outside the driving domain. To this end, we introduce DriveStreamBench, a driving and streaming video benchmark that organizes this capability spectrum into four cognitive levels constructed from a shared pool of driving events, to support level-wise capability profiling and gated evaluation. We evaluate various vision-language models along with DriveStream-SFT, a supervised reference baseline trained on the benchmark's training split, under a unified streaming protocol. DriveStreamBench tests a basic but necessary capability, and all evaluated models fall short of it; current vision-language models are not yet ready for user-conditioned alerting in driving. Resources are available at https://anonymous.4open.science/r/DriveStreamBench/ and https://dataverse.harvard.edu/previewurl.xhtml?token=01f6bc4c-d809-4e40-b025-b6d783c86571.