Do Neural Foundation Models Help Yet? Comparison of Self-Supervised Objectives for Intracortical Decoding
Sara Cammarota ⋅ Tommaso Boccato ⋅ Nicola Toschi ⋅ Matteo Ferrante
Abstract
Inspired by foundation models in vision and language, recent work has begun to explore large-scale self-supervised pretraining for intracortical neural decoding. However, whether the scale of data currently available is enough to see scaling benefits emerge remains contested, and existing studies differ simultaneously in objective, architecture, corpus, and downstream task. We present the first systematic comparison of self-supervised paradigms for invasive decoding in a unified framework. Holding a common transformer backbone fixed, we compare masked autoencoding (MAE), Joint-Embedding Predictive Architectures (JEPA), autoregressive (AR) pretraining and training from scratch on up to $345$ hours of recordings: $208$ hours from $6$ human participants performing attempted and imagined speech, handwriting, typing and cursor control, plus $137$ hours of motor tasks from $19$ non-human primates (NHP). We find that the pretraining gains are often comparable to carefully tuned task-specific training. These findings establish a controlled benchmark for future research on neural foundation models.
Chat is not available.
Successful Page Load