Next-Token Prediction Enables Scalable Learning of Sleep Physiology
Abstract
Foundation models offer a promising route to compress multi-modal physiological signals into better measures of human health, with broad applications across sleep medicine, cardiology, neurology and other healthcare domains. Existing models are typically trained with contrastive or masked-reconstruction objectives. However, masked reconstruction may be poorly suited to the stochastic nature of these signals, while contrastive approaches rely on positive-pair definitions despite the semantic invariances of physiological signals being poorly understood. In this work, we show that next-token prediction is a simple and scalable alternative. We develop Hypnos, a multi-modal sleep foundation model trained via next-token prediction using eight different sensing modalities (e.g. EEG, ECG, respiratory signals) from overnight polysomnography recordings. We tokenize each modality into streams of discrete tokens using residual vector quantization. We then train a large auto-regressive RQ-Transformer to jointly predict the next token across all modalities in parallel, using a novel modality masking strategy to enable generalisation to subsets of modalities during inference. Using over 20,000 overnight recordings drawn from nine public datasets, we find that both next-token perplexity and downstream probing performance continue to improve with model scale. Across a range of tasks, Hypnos matches or exceeds prior sleep foundation models and strong supervised baselines on sleep stage classification across in-domain and held-out test sets. Our results indicate that next-token prediction is a strong self-supervised objective for learning representations from multi-modal physiological signals.