From Detection to Understanding — A Multi-Task Dataset for Traffic Anomaly Reasoning
Abstract
We present TAR (Traffic Anomaly Reasoning), a large-scale multi-task dataset for training and evaluating vision-language models on traffic anomaly understanding. While video anomaly detection (VAD) has traditionally focused on binary classification or temporal localization, TAR moves beyond detection to reasoning. TAR contains 44,040 training annotations with explicit chain-of-thought reasoning traces across 10 task types, covering 3,670 CCTV transportation videos (~26 hours) sourced from eight public datasets. TAR provides a structured progression of tasks organized into three groups — Question Answering, Temporal Reasoning, and Scene Understanding — reflecting our finding that reliable anomaly understanding demands capabilities far beyond detection. Annotations are produced by a hierarchical auto-labeling pipeline using state-of-the-art VLMs, with existing human annotations incorporated as supplementary context for 25% of videos. Every answer is paired with a step-by-step reasoning trace, enabling both supervised training with chain-of-thought supervision and fine-grained evaluation of model reasoning. We evaluate eight vision-language models on a held-out test set of 84 videos (1,008 annotations) with human-reviewed and corrected annotations and find that current models struggle with Temporal Reasoning and Scene Understanding tasks even when they perform well on basic Question Answering — revealing that detection accuracy is a poor proxy for genuine understanding. A multi-task ablation study shows that progressively adding task groups during fine-tuning yields consistent gains, with the full 10-task model achieving +44.5 points on Binary QA accuracy and +21.4 points in aggregate score over the zero-shot baseline, demonstrating the dataset's value as a training resource. TAR is released under CC-BY-4.0 and serves as the official training data for AI City Challenge 2026 Track 3: Anomalous Events in Transportation. The annotations and video download scripts are publicly available at https://huggingface.co/datasets/nvidia/PhysicalAI-Traffic-Anomaly-Reasoning.