State-Transition Supervision for Reliable Temporal Video Models
Abstract
Temporal video models often learn dynamics only from appearance and labels. We study a complementary signal: an explicitly estimated state whose allowed transitions are known. In a controlled speedcubing task, a model predicts one of six moves from a trimmed RGB clip, while a separate pipeline estimates the visible 3×3 cube state. We regularize the move distribution according to whether observed state changes are compatible with each candidate move. The state signal is used only during training. Across ResNet3D, ResNetLSTM, ResNetGRU, and ResNetTransformer models, the best state-aware settings improve F1 by 0.017–0.058 over matched visual-only baselines; ResNet3D reaches 0.828 F1. The gain disappears when the auxiliary term is overweighted, and stress tests identify occlusion as the main failure mode. This compact setting isolates a design question relevant to larger temporal systems: explicit state dynamics can improve reliability, but only when the state estimate is treated as uncertain auxiliary evidence rather than ground truth.