CREF: Forecasting Benchmarks for the Age of Agents
Abstract
Traditional forecasting evaluations relying on static historical data are fundamentally incompatible with modern AI advances. Even for standard time series foundation models, temporal overlap between training and test periods can lead to implicit leakage through correlated signals. This problem is intensified by LLMs, which can bridge semantic gaps between forecasted and related data while memorizing extensive world knowledge. At the extreme, agentic models with tool access can retrieve the ground truth directly. Such information leakage inflates performance relative to real forecasting scenarios and is impossible to rule out in historical evaluations. We introduce CREF, a real-time benchmark that evaluates forecasting models on data that does not yet exist, eliminating all forms of leakage by design. CREF balances three competing objectives for real-time benchmarks: diversity, robustness, and throughput. It draws data from 32 public APIs and features 50 forecasting tasks with diverse data frequencies, multivariate targets, and covariates, making it the first real-time benchmark reaching the breadth of leading static benchmarks. Retrospective evaluation validates that CREF produces rankings consistent with established static benchmarks, while real-time evaluation reveals qualitative and quantitative performance gaps for agentic and LLM-based models, confirming leakage on historical data. CREF runs continuously and is open for new models to join, offering a public leaderboard for contamination-free model comparison.