Knowing When Multivariate Forecasts Are Wrong
Abstract
At deployment, a multivariate forecaster may expose only a recent input window and a frozen point prediction, with no access to ground truth, calibration residuals, repeated inference, or model internals. We study deployment-constrained forecast risk localization: ranking future timesteps by likely error before they are observed. We introduce SubDx, a zero-parameter diagnostic that audits the forecast against the local cross-channel subspace of the input history. Each predicted channel vector receives an off-subspace residual score, computed with no labels, no learned parameters, and less than 1% runtime overhead on Traffic. Under a linear factor model, we derive closed-form score distributions and a two-population AUROC expression showing how heterogeneous off-subspace error visibility makes the ranking identifiable. On Traffic (N=862), SubDx reaches within-horizon AUROC 0.90 and reduces MSE by 62% at 50% coverage; on Solar (N=137), it reaches AUROC 0.85 with 67% MSE reduction. The same post-hoc score applies to frozen pretrained forecasters, with Chronos reaching AUROC 0.838 on Traffic without fine-tuning. Across trained backbones, pretrained models, and synthetic controls, the results show that local cross-channel consistency can provide a practical abstention signal for redundant multivariate systems.