What If TSF: A Benchmark for Reframing Forecasting as Scenario-Guided Multimodal Forecasting
Abstract
In multimodal time series forecasting, predictive accuracy alone provides limited insight into how forecasts respond to different external scenarios. We introduce What If TSF (WIT), a benchmark for evaluating fine-grained forecast responsiveness to future scenarios. WIT constructs controlled scenario sets that hold the forecasting state fixed while varying factual, domain-grounded alternative, and irrelevant scenarios. It assesses scenario responsiveness along three complementary dimensions: Factual Comparison measures the predictive utility of factual scenario information; Relational Comparison evaluates whether forecasts satisfy expected direction, contrast, intensity, and restraint relations across scenarios; and Grounded Comparison assesses whether predicted response trajectories align with those observed in matched real cases. Experiments on WIT show that factual accuracy gains do not consistently translate into appropriate responses across alternative scenarios, highlighting a distinction between predictive accuracy and scenario responsiveness.