Can VLMs Keep Time? Probing Quantitative Temporal Reasoning in Physiological Videos
Abstract
Vision-language models (VLMs) are increasingly used for multimodal medical reasoning, yet it remains unclear whether they can reliably recover and reason over the temporal structure underlying physiological measurements. We study this question through contactless respiratory monitoring and introduce a paired temporal abstraction ladder that evaluates the same physiological episode as raw video, temporal signal, explicit respiratory events, and symbolic temporal facts. This design separates failures in visual extraction, temporal abstraction, and downstream quantitative reasoning rather than collapsing them into a single end-to-end error. Across three respiratory-video datasets, four reasoning tasks, and four frozen VLM backbones, macro-accuracy rises from 50.6% at the raw-video level to 90.9% when temporal facts are explicitly provided, revealing a 40.3-point L0-L3 abstraction gap. Paired transition analysis further shows that progressively exposing temporal structure rescues 43.2-61.7% of preceding errors across successive abstraction levels. These results suggest that current VLMs often possess the downstream reasoning capability required by the evaluated tasks once temporal structure is explicit, but remain substantially less reliable at deriving that structure directly from video.