Scaling Neural Motor Decoding via Decoupled Behavioral Pretraining
Abstract
Recent neural decoding approaches achieve strong performance by training on large collections of paired neural-behavioral recordings, but such datasets are expensive, invasive, and difficult to scale. In contrast, behavioral data is abundant and easy to collect, raising the question of whether neural decoding can benefit from decoupling behavioral and neural representation learning. We introduce BeeMO, a flexible multi-session framework that enables training with arbitrary mixtures of paired neural-behavioral recordings and unpaired behavioral data. A behavior decoder learns movement dynamics directly from behavioral trajectories, while neural activity is incorporated through cross-attention layers when available, allowing behavioral and neural supervision to scale independently. By decoupling neural and behavioral data, we unlock a new dimension for scaling: unpaired behavioral data, which is substantially cheaper and easier to collect than paired neural recordings, can be used directly to improve decoding performance. Across multiple intracortical motor datasets and tasks in nonhuman primates, we show that incorporating unpaired behavioral data consistently improves decoding performance, particularly, and that incorporating task-aligned behavioral trajectories can further improve transfer. We further show that behavior-only pretraining, without any neural pretraining, outperforms single-session supervised baselines and is comparable to models pretrained on substantially larger neural datasets. Together, our results suggest that scaling behavioral data offers a practical and cost-effective path toward neural decoding models that generalizes across subjects and tasks.