WhatIf-Bench: Benchmarking Intervention-Aware Time Series Forecasting
Arjun Ashok ⋅ Mihir Parmar ⋅ Palash Goyal ⋅ Bhavana Dalvi Mishra ⋅ Chun-Liang Li ⋅ Irina Rish ⋅ Alexandre Drouin ⋅ Jinsung Yoon ⋅ Tomas Pfister
Abstract
Recent time series foundation models (TSFMs) can consume covariates in-context, making them applicable to real-world forecasting driven by external factors, and to scenario analysis: forecasting how a target would respond to events that set its covariates to chosen future values. Whether they can do this reliably, however, is still an open question. We introduce WhatIf-Bench, a benchmark that measures how well a forecaster can learn covariate--target relationships in-context and extrapolate them to the future. We adapt mechanistic simulators of real-world systems, spanning data centers, power grids, road traffic, and water networks, so that any covariate can be intervened on and its effect on a target observed. We generate $40,000$ tasks whose future event either replays seen covariate values (a Replay task) or sets them to unseen values (a Novel task), varying the number of covariates and the number and length of historical events. Evaluating a range of covariate-aware TSFMs, we find that performance varies widely: TimesFM-3 and TS-ICL reach the highest skill scores ($67\%$), while the remaining models cluster lower. Every model forecasts well on Replay futures ($60$--$85\%$ skill) but worsens significantly on Novel ones (best $48.6\%$), and falls further as events set more covariates jointly, though a longer history of events partly recovers this. WhatIf-Bench reveals gaps in TSFMs that restrict their reliability for arbitrary scenario analysis, and offers a testbed toward models that can plan across possible futures. WhatIf-Bench will be open-sourced upon publication.
Chat is not available.
Successful Page Load