Disentangling Where and When: Factored Spatio-Temporal Explanations for Video Action Recognition
Bhavan Vasu ⋅ Giuseppe Raffa ⋅ Prasad Tadepalli
Abstract
Post-hoc explanation methods for video action recognition typically conflate spatial appearance and temporal dynamics into a single saliency or importance map. We propose a perturbation-based framework that disentangles these contributions into a spatial appearance mask ($M_a$) and a temporal motion mask ($M_m$). The two masks are jointly optimized against a frozen action recognition model through factored perturbation architecture, with nearest frame replacement to identify important temporal dynamics. A separation penalty discourages the appearance mask from spending budget on frames that the motion mask has already discarded, and a one-sided ReLU area constraint allows the realized masks to fall below the optimization budget when the input does not require its full extent. To evaluate these explanations, we introduce a separation protocol that sweeps spatial pixels and temporal frames independently, backed by proofs characterizing the inherent geometric bias of standard Insertion metrics. Across EPIC-Kitchens-55, EGTEA Gaze+, and Something-Something V2, our method attains the lowest deletion AUC against leading baselines while producing masks roughly five times sparser.
Chat is not available.
Successful Page Load