Pixel-level understanding of a world in motion within a neural encoding framework
Abstract
In this work, we provide a holistic study of neural encoding using fMRI recordings of the human brain for participants watching video stimuli and how it relates to deep neural networks. Our work advances beyond recent studies that is focused on coarse image- or video-level object and action recognition models in neural encoding. We instead study deep networks designed for pixel-level understanding tasks, e.g., motion and depth estimation and how they impact the neural encoding in comparison to coarse models, e.g., action or object recognition models. Moreover, we use this to study class-agnostic vs. class-aware tasks within neural encoding. The architectures studied include both convolutional (e.g., ResNets) and transformer-based models (e.g., ViTs) to ensure a comprehensive study. Additionally, this study covers both the dorsal and ventral visual regions within a neural encoding framework. We conduct extensive experiments on both the Mini-Algonauts and the BOLD moments datasets studying 20 models for object recognition, action recognition, optical flow estimation, depth estimation and semantic segmentation. We learn a regularized regression for neural encoding using a technique that takes the hierarchical representations of a deep network from multiple layers as input to predict the voxels of a single region of interest. We use the Pearson’s correlation coefficient between the actual and predicted voxels’ responses as our metric, its noise-normalized variant and study the layers’ contribution patterns.