A Structural Load Model Cannot Predict the Effect of a Conservation Appeal
Jayvant Rajesh ⋅ Wenhao Lu
Abstract
Temporal foundation models are increasingly asked to act as world models, imagining alternative futures under interventions rather than only forecasting. We ask what happens when a structural temporal model built for that purpose is held to the stronger standard. We decompose CAISO system load over 2016 to 2024 into a thermal layer, hour-of-day behavioral coefficients identified by matched difference-in-differences, and a learned residual, then subject it to four holdouts. On the 2023 to 2024 accuracy holdout, coefficients from a careful matched design and from naive full-series regression differ by 0.04 MAPE points, so the forecasting metric cannot rank the two designs. On an interventional holdout of the ten consecutive Flex Alert days of the September 2022 California heat wave, every model predicts a nearly constant appeal effect of one to two percentage points while the observed effect ranges from $+7.4$ to $-19.9$; a baseline predicting no effect attains the lowest error, and the identified model's per-day ranking is significantly anti-correlated with the truth ($\rho_s = -0.76$, $p = 0.011$). A held-out lockdown regime is the counterweight: removing the occupancy term costs 4.9 MAPE points, so structure carries information where identification does not. The shift that defeats the model is in the treatment rather than the covariates, so leakage-aware splits and contamination audits cannot see it. We propose the interventional holdout as an evaluation primitive for temporal models with world-model ambitions, with a mandatory zero-effect comparator, and report a serving-time trust region that flags the failure without ground truth.
Chat is not available.
Successful Page Load