Is Temporal Order Load-Bearing? A Shuffle-and-Blind Audit of Multi-Frame Driving VQA
Abstract
Multi-frame benchmarks are increasingly used to claim that vision-language models (VLMs) perform temporal reasoning in driving scenes. We argue a benchmark supports such claims only if frame order is load-bearing, and we propose a three-arm audit that any multi-frame benchmark can run at negligible cost: ordered frames, shuffled frames (seeded permutation while the prompt still asserts temporal order), and blind (no images). Applying the audit to the publicly released Automingo-VQA dataset (5 frames spanning +/- 2 s, n = 1,055 validation questions) with three open VLMs spanning model families, architectures, and serving stacks - Gemma-4-31B (dense), Gemma-4-26B-A4B (sparse MoE), and Qwen3-VL-30B-A3B (a non-Gemma MoE from the same generation as the benchmark authors' fine-tuning base) - we find frame order contributes nothing in any of them: 83.32% ordered vs. 84.08% shuffled for the 31B (90 discordant pairs, McNemar p = 0.461, shuffled nominally higher), identical 81.14% vs. 81.14% with exactly symmetric 30:30 discordance (p = 1.0) for the Gemma MoE, and 80.47% vs. 80.66% (24:26, p = 0.888) for Qwen. Prompted chain-of-thought reasoning does not change the verdict in either family (p = 0.915 MoE, p = 0.657 dense) - and on the dense model it significantly reduces accuracy (p = 0.0005) by collapsing precisely the motion-defined situations: reasoning tokens neither use nor recover temporal order, and the timelines they confabulate carry a real cost. Score decomposition attributes the benchmark entirely to a text prior (66.9% blind) plus single-scene static evidence (+19.2 points), leaving +0.0 for temporal order - and the null holds within motion-defined situations (cut-in, merging, leading-braking), whose per-situation deltas scatter in both directions. Both audited models score in the frontier band on this benchmark, so the null is not a weak-model artifact - and its replication across a dense and a sparse-MoE architecture on two different serving stacks indicates a property of the benchmark, not of one model: systems matching fine-tuned and proprietary baselines treat the sequence as a bag of frames, exercising no inference-time temporal world model. We recommend that multi-frame embodied benchmarks report shuffle and blind arms as standard practice, and publish our audit protocol details, per-item pairing convention, and full run provenance.