CorticAll: A Universal Cortical Decoder
Abstract
Intracortical brain--computer interfaces (BCIs) for communication resist the data-scaling that transformed neighbouring automatic speech recognition (ASR), because their input is heterogeneous: noisy, multivariate, permutation-sensitive, and of participant-specific dimensionality. Decoders are therefore trained one participant at a time, forgoing the benefits of pooled data. We introduce CorticAll, a multimodal cortical decoder that ingests intracortical signals of arbitrary channel dimensionality and decodes speech, handwriting, typing, and cursor control within a single model. A shared per-channel front-end tokenizes each channel, and a cross-attention Channel Mixer collapses an arbitrary number of channels to a fixed-size token, enabling fully end-to-end training over a heterogeneous pool of 11 datasets from 7 participants spanning 4 cortical regions and 4 modalities (approx. 173h). We further introduce supervised-on-scale pretraining---joint supervised training across all datasets and modalities---and argue it is a more data-efficient alternative to self-supervised pretraining. Fine-tuning the pooled model reaches state-of-the-art error rates on 7 of the 8 datasets with published baselines and decodes cursor velocity at up to R^2=0.93. Analysing the learned participant x modality latent space, we find that speech, handwriting, and typing reuse a largely shared subspace---evidence of reusable neural primitives across modalities.