ColdCast: A Benchmark and Evaluation of Foundation Models for Cold-Start Forecasting
Abstract
Cold-start forecasting (CSF), predicting targets for new entities with no historical observations, is an open challenge actively being researched in the industry, yet largely overlooked in the forecasting literature. A key barrier to progress is the absence of standardized evaluation: to date, no dedicated benchmark exists for CSF. We address this gap by introducing ColdCast, a cold-start forecasting benchmark spanning five domains and 140,000 series including news, books, games, fashion, and movies, enabling systematic evaluation under genuine cold-start conditions. Using ColdCast, we conduct an empirical study of zero-shot Time Series Foundation Models (TSFMs) and Prior-Fitted Networks (PFNs) - two distinct classes of zero-shot forecasters - and baselines including population-level statistics computed from existing entities. We find that PFNs (TabPFN-TS and ApolloPFN) outperform TSFMs, and are the best-performing model on 2 of 5 datasets. Simple population-level averages are a strong baseline, achieving the best performance on the remaining 3 datasets. For PFNs, we show that increasing the number of in-context analogous entities, and incorporating metadata-derived embedding features each yield additional gains in the zero-shot setting.