RABBiT: Interactive speech-to-fMRI prediction and in-silico fMRI experiments
Abstract
RABBiT is a real-time speech to fMRI foundation model; it lets attendees turn their speech into brain predictions in real-time, and it also enables running neuroscience experiments in real-time. Enabled by efficient and strong design, RABBiT predicts accurate zero-shot population-level fMRI responses across auditory and language regions. The demonstration combines an interactive cortical viewer for real-time speech-to-fMRI prediction with the ability to also run classic experiments like the classic language-localizer task to explore how predicted responses contrast when speech is intelligible vs when it is not.
RABBiT was pre-trained on a large paired naturalistic audio-fMRI dataset, making it outperform the state-of-the-art on zero-shot fMRI prediction accuracy while being light enough to run on browser real-time.
We enable attendees to try three activities:
- Explore speech-evoked responses: speak into a microphone and watch predictions update as audio is processed real-time. Rotate the brain to explore the response pattern.
- Run a guided fMRI experiment: choose individual pairs of contrasts (e.g., intelligible vs unintelligible speech ), switch between condition and difference maps, and inspect regional values.
- Mimic an fMRI experiment: read a supplied text like you would in an fMRI experiment, then the contrast, and both conditions run through RABBiT real-time. Then, one can inspect their predicted response difference and download the results. Audio is automatically not retained for privacy reasons.
We think this is a step forward in the direction of democratizing access to in-silico neuroscience, where people get to try classic experiments and devise new ones and compare results, without the need to be an expert on collecting the data. It also fosters curiosity and easy access to testing initial ideas before committing to collecting human data. We hope that this step will be followed by feedback from users to improve this kind of foundation models and extend it to other imagining modalities (e..g, EEG and MEG).
The demo connects the workshop’s focus on foundation models for neural data with reproducible neuroscience experiments, giving attendees a practical way to probe model behavior and develop hypotheses for subsequent fMRI studies. We commit to making it public for everyone to use by the time of the workshop so that this effort can commence right away.