Evaluating Neural Data Tokenizers: A Framework for Assessing Learned Representations of Spiking Activity
Abstract
Foundation models of brain activity aim to learn representations from large-scale neural recordings that generalize across sessions, subjects, species, and downstream tasks. Transformer-based models have emerged as a promising architecture for this goal and begin all with the same first step --- tokenization: transforming sparse binary events into vector representations that downstream networks can consume. In other areas of machine learning, tokenization is its own design problem, distinct from the downstream foundation model. Several mostly implicit tokenization approaches dominate the literature, but their impact on downstream performance has not been systematically evaluated. We propose four explicit desiderata for a neural tokenizer---reconstruction fidelity, downstream utility, cross-instance generalization, and compression---and a framework of metrics spanning seven dimensions, including firing statistics, encoding and decoding, and latent structure and robustness. We assemble a benchmark of ten datasets spanning species, brain areas, recording technologies, and behavioral paradigms, and evaluate a representative set of discrete and continuous tokenizers, both single-neuron and population variants. Across the strategies evaluated, we show how tokenizers trained only on reconstruction can support downstream tasks. Yet reconstruction alone is misleading: reconstruction quality and downstream utility can be sharply dissociated. Our multi-axis framework is a response to this failure mode. These findings highlight tokenization as a critical and underexplored design axis in brain foundation models, and suggest that further progress will require continued systematic study of tokenization strategies.