BrainWhisperer: Leveraging Whisper for Speech Decoding in Neuroprosthetics
Abstract
Decoding continuous speech from intracortical recordings holds transformative potential for individuals with severe motor impairments, yet current systems face fundamental barriers to real-world deployment: training data is scarce, neural signals are non-stationary across sessions, and the external language models required by state-of-the-art cascaded decoders impose memory requirements that preclude local, privacy-preserving inference. We introduce BrainWhisperer, a neural speech decoder that adapts Whisper–pretrained on approximately 680,000 hours of speech–to microelectrode array recordings via a convolutional front-end, hierarchical low-rank projections for non-stationarity, windowed self-attention in phoneme-selective encoder layers, and a multi-task objective combining CTC and cross-entropy losses. Subject-specific embedders enable cross-participant training within a unified architecture. Evaluated on publicly available Utah array datasets from BrainGate participants, BrainWhisperer achieves a word error rate of 8.5\% in end-to-end decoding on the Card benchmark--to the best of our knowledge, the best result reported in this setting. Cross-dataset training improves performance without participant-specific fine-tuning, pointing toward scalable foundation models for speech brain-computer interfaces.