Intervenable Concept Dynamics for Video Activity Forecasting
Abstract
Concept bottleneck models route predictions through human-understandable concepts, making their intermediate states inspectable and editable. Extending this model family to forecasting requires modeling evolving concept activations, yet existing CBMs ground concepts only in observed inputs and do not forecast their future evolution. We propose TRACE-CBM (Temporal Relational Activity Concept Evolution) for video activity forecasting through a named-concept state: it refines observed concept activations with sparse spatio-temporal graphs, recursively predicts future states conditioned on prior activity distributions, and decodes both with a shared activity classifier. At inference time, the model supports editing concept values, activity activations, and learned edges without parameter updates, exposing the relational spatio-temporal dependencies used for forecasting. Across multiple datasets, TRACE improves observed-activity Top-1 accuracy over concept-based methods by 1.2--8.5 percentage points and attains H3 forecast Top-1 accuracy within 1.9--2.2 points of matched black-box forecasters. Intervention experiments show concept-node edits correct 25.6--34.4% of initially incorrect H3 forecasts without flipping initially correct ones, and concept steering achieves the requested target activity in up to 90.0% of cases.