From Raw Recordings to Model-Ready Batches: Unifying BCI Datasets for Foundation Models
Abstract
demo link :Speech brain-computer interfaces (BCIs) are scaling past single-participant proofs-of-concept. Multiple input modalities (e.g., Utah arrays, flexible threads, endovascular stentrodes, ECoG grids, sEEG, EEG, MEG, fNIRS, fMRI) now interface with decoders emitting text, synthesized voice, or avatar control. Progress is gated by data scarcity and lack of standardization: ECoG studies typically include <15 participants and minutes-to-hours of recording; intracortical studies involve few intensively trained implanted users; semantic fMRI requires hours per participant. Channels don’t correspond naturally across individuals, and naive pooling may erase meaningful distinctions. Models need thousands of labelled trials per participant and transferable shared representations. Closing this gap requires shifting the data infrastructure BCI models are built on.
Many high-value neural datasets exist which could be used to train BCI, however going from source data files to a model-ready batch, temporally aligned with rich task labels, requires substantial processing. NeuralSet is a Python framework developed for this purpose that decouples lightweight event metadata from lazy tensor extraction. This separation allows studies to be filtered, recombined, and batched as PyTorch datasets before source data is loaded, while tensor extraction can be deterministically cached and dispatched across compute clusters. However, it is a very recent, general framework that lacks some of the specific, frequently needed preparatory steps.
We propose a demo of our extension of NeuralSet to ingest open BCI datasets, featuring a speech BCI model. We have built a set of onboarding recipes to ingest relevant public datasets by mapping their structures (file organization, signal streams, event annotations, linguistic labels, identifiers, channels, sampling conventions and temporal references) onto a shared metadata schema while preserving modality and study-specific information. datasets are searchable, temporally aligned, and accessible via a common interface, allowing heterogeneous recordings to be readily pooled into model-ready batches for training neural foundation models. In the demo, attendees will run Jupyter notebooks that read these datasets directly into NeuralSet, and interactively load, explore and visualize large-scale recordings. These datasets were used to train BrainWhisperer, a lightweight speech BCI model built on this stack. BrainWhisperer adapts the Whisper model to microelectrode array recordings via subject-specific embedders and hierarchical month/day projections, achieving ~8.5% word error rate in end-to-end single-dataset decoding and better performance under cross-dataset training without participant-specific fine-tuning. As training is time-prohibitive, the demo focuses on inference: attendees will decode sentences from real recordings with the trained model and inspect how unified datasets across sessions, participants, and devices were concatenated for joint training.
For those building foundation models over neural data, the demo is a working illustration of how to turn raw recordings into model-ready batches. It shows researchers how to harmonize raw recordings and the thousands of labelled trials such models require into a unified standard without having to manually process the data themselves.