Direct Conditioning of Audio Diffusion Transformers on fMRI Reveals Cortical Contributions to Sound Reconstruction
Abstract
Reconstructing audio from brain activity remains challenging. A key limitation lies in current audio diffusion models, which are primarily designed for text-based conditioning and lack mechanisms to incorporate non-linguistic continuous signals such as neural responses. While recent work in the visual domain has explored direct conditioning on brain activity, this approach has not yet been extended to audio diffusion models. We introduce a framework that directly guides a latent audio diffusion model with fMRI signals, avoiding intermediate feature prediction for the conditioning stage. Central to our approach is a transformer-compatible adapter that enables the integration of non-linguistic representations into the cross-attention layers of diffusion transformers (DiTs). Although motivated by neural decoding, this adapter provides a general mechanism for conditioning audio diffusion models on arbitrary continuous inputs, not inherently limited to neural signals. We demonstrate consistent improvements over baselines, including two-stage fMRI-to-latent conditioning, across both acoustic and semantic evaluation metrics. In addition, we introduce a fidelity metric grounded in a computational model of auditory neural processing, which allows us to quantify the contribution of individual auditory cortical regions to reconstructed audio representations.