Benchmarking and Personalizing Time-Series Foundation Models for Continuous Glucose Forecasting
Abstract
Continuous glucose monitoring (CGM) produces dense, patient-specific time series, but it remains unclear whether general-purpose time-series foundation models (TSFMs) can effectively forecast glucose across clinically relevant horizons or benefit from wearable covariates and personalization. We benchmark Chronos-2 and TimesFM 2.5 against CGM-specific models and conventional baselines on AI-READI, a multi-site cohort spanning the spectrum of diabetes severity, across forecast horizons ranging from 30 min to 4 h. Both TSFMs outperform persistence, linear extrapolation, and circadian baselines across all horizons, and outperform both CGM-specific models evaluated zero-shot; of those, CGM-LSM also exceeds it. An LSTM and Mamba trained from scratch perform comparably to TimesFM, suggesting that pretraining scale alone does not explain TSFM performance. Additional covariates provide modest, model-dependent benefits and some measurable harms: several wearable channels improve Chronos-2 at longer horizons only, while time-of-day encoding degrades it at every horizon and predicted meal probability degrades TimesFM at 30 min to 2 h. Personalization yields the largest gains, outperforming LoRA fine-tuning with the same history length. These findings identify an individual's recent history, rather than model scale or additional wearable channels, as the most effective source of adaptation for CGM forecasting.