Emotion-Trained Vision Models Do Not Necessarily Learn EEG-Aligned Facial Dynamics
Abstract
Facial expressions are one of the main ways humans communicate affect, but they are not static signals. In natural vision, expressions evolve over time, with facial movements emerging, intensifying, and changing as emotion is expressed. Prior model-brain studies based on static images suggest that vision models can capture aspects of face-related neural representations, but it remains unclear whether artificial vision models capture the temporal EEG dynamics of facial emotion perception. We test the hypothesis that models trained for facial emotion recognition, especially video-based or temporally structured models, should show stronger alignment with human EEG dynamics than randomly initialized controls. We do so by aligning model representations with time-resolved EEG representational geometry from 25 participants viewing one-second dynamic facial-expression videos. We compare six models spanning static-image, video-based, self-supervised, and vision-language representations: ResNet18, VGGFace, CNN2D, CNN3D, DINOv2-temporal, and Qwen2.5-VL. Contrary to this hypothesis, supervised facial-emotion training does not consistently increase alignment with human EEG dynamics. In particular, video-trained and temporally structured emotion models align less with fear-related EEG geometry than the same architectures with randomly initialized weights, whereas static-image emotion models show trained-versus-random differences close to zero. These results are robust to stricter noise-ceiling thresholds and are not explained by actor identity or raw pixel similarity. Exploratory analyses further suggest that peak alignment is often localized to the first or last sampled video frames, and that fear alignment varies across scalp ROIs by training regime. Overall, our results challenge the assumption that video-based emotion supervision automatically yields more brain-like dynamic representations of facial emotion.