CrashJepa: Auditing Temporal Grounding in V-JEPA Accident Anticipation
Abstract
Temporal foundation models can make safety-critical video tasks appear solved under recognition metrics, while still relying on order-insensitive context. We study this failure mode in dashcam accident anticipation using frozen V-JEPA representations. V-JEPA-only and V-JEPA+trajectory baselines obtain strong accident recall, but often preserve recall when temporal order is shuffled. We therefore frame accident anticipation as a temporal-grounding problem rather than only a video-classification problem. We introduce CrashJepa, a multi-head V-JEPA and trajectory-interaction model with a grounded urgency-aware temporal (GUT) loss that encourages warning-score changes to follow the direction of a trajectory-derived danger proxy. We evaluate chronological input against full temporal shuffle, reversal, block-order shuffle, and feature-level object masking on CCD, DAD, and Nexar. With GUT, recall drops sharply under destroyed temporal order, e.g. CCD from 99.9% to 15.7% under full shuffling and Nexar from 84.9% to 6.6% under reversal. These results position CrashJepa as a case study in reliability evaluation for temporal foundation-model systems: fixed-threshold recognition results should be audited for sensitivity to temporal ordering and evidence from road-agent interactions.