TRACE: Trajectory-Aware Conceps for Explainable Video Understanding
Abstract
Video understanding models often rely on motion, but faithful explanations must reveal not only whether motion matters, but also which source of motion supports a prediction. We identify motion-source under-specification as a key limitation of existing video XAI: region-centric explanations are source-ambiguous, while concept-based explanations are source-incomplete when their motion vocabulary is restricted to single-person pose sequences. To bridge this gap, we propose TraCE-Trajectory-Aware Concepts for Explainable Video Understanding-a concept-based framework that elevates visual trajectories to first-class explanatory concepts. By clustering optical-flow trajectories, TraCE discovers source-flexible motion concepts that capture dynamic evidence from people, objects, and interactions, while object and scene concepts account for static context. TraCE applies the same typed trajectory/object/scene concept interface to both action recognition and video question answering (VQA), enabling concept-level explanations beyond a single task. We evaluate TraCE across six benchmarks: NTU RGB+D Mutual, Chiral SSv2, UCF-101, KTH, TOMATO, and STAR. Experiments show that TraCE provides trajectory-aware explanations of dynamic evidence with minimal performance trade-off, trailing non-interpretable baselines by only 0.9 points on action recognition and 0.1 points on VQA on average. User studies, faithfulness analyses, temporal sensitivity tests, and concept-intervention results further show that TraCE produces interpretable, faithful, and practically useful explanations.