BRIDGE: Brain-Vision Representation Integration through Depth and Granularity Encoding
Abstract
Decoding visual content from non-invasive brain signals remains challenging because neural responses evolve over time whereas images are typically static. Most existing methods align an entire neural response window to a single final visual embedding, this overlook a fundamental representational mismatch: brain signals reflect both early perceptual and later integrative representations, whereas the final visual embedding is biased toward high-level semantics. Therefore, we propose BRIDGE, a brain-vision representation integration framework that aligns two modalities through Visual Depth Encoding and Brain Granularity Encoding. On the visual side, BRIDGE extracts and fuses CLIP representations from multiple depths, producing an alignment target that preserves low-level and high-level information. On the brain side, instead of treating the whole temporal window as homogeneous and obscuring temporally heterogeneity, BRIDGE explicitly partitions stimulus-evoked EEG/MEG responses into a small number of temporally ordered stages and adaptively pools them. The resulting brain and visual embeddings are trained in a shared latent space by contrastive learning and can further support brain-to-image generation through a pretrained diffusion prior. Experiments on THINGS-EEG and THINGS-MEG demonstrate that BRIDGE achieves strong retrieval and generation performance. Ablation studies further confirm the complementary benefits of depth-wise visual aggregation and temporal brain factorization.