V1-Inspired Dynamic Vision System: A Bio-Plausible Video Embedding Framework with Decoupled Shape–Color Pathways and Long-Range Spatiotemporal Perception
Abstract
Large-scale video–text contrastive learning has become the de facto standard for video understanding, yet mainstream methods almost universally adopt the linear patch embedding of vision Transformers (ViT) or evenly sliced tubelets as the front end. We revisit this front end from the biological principles of the primary visual cortex (area V1) and introduce V1DynamicVisionSystem (V1-DVS), a fully end-to-end trainable, bio-plausible video embedding framework. Its main contributions are: (i) a three-pathway parallel architecture that maps onto layer 4C orientation selectivity, layer 4B motion direction selectivity, and the cytochrome-oxidase Blob color-opponent system; (ii) an explicit shape–color decoupling, with isotropic Blob processing running in parallel with anisotropic Gabor orientation processing; (iii) long-range spatiotemporal convolution that captures inter-frame optical flow through 3D filters and the Adelson–Bergen motion-energy model rather than naive frame stacking; and (iv) bidirectional cross-pathway attention, allowing motion and color features to mutually modulate, as in V1→V4 inter-areal feedback. Pretrained on CC3M, WebVid-10M, ActivityNet and Ego4D and evaluated under a fully self-supervised objective(SimCLR + temporal-order verification + masked-token prediction), V1-DVS outperforms four strong baselines—linear patch embedding, mean temporal pooling, tubelet embedding, and spatiotemporal latent encoding—on six metrics spanning retrieval and embedding-space geometry, validating the practical value of biologically grounded structure for video embedding.