When the Stratifier Is Not Identified: A Pre-Registered Self-Audit of Regime-Stratified Evaluation
Abstract
Regime-stratified evaluations report that a model fails in a named condition, and the number is read as a statement about that condition. It is only that if the label carries information the forecaster did not already have. We audit our own attempt to measure this on a frozen panel of time-series foundation models, under a pre-registration that fixed the estimand, the decision rule, and the falsification controls before the runs. The audit refutes our own headline. Replacing the labels with a null that destroys their meaning but preserves their structure reproduces the reported gap in full: 0.364781 and 0.400192 mph against a reported 0.350113 mph, both null intervals excluding zero, with every one of 200 shuffles positive in both constructions. The pre-registered reading of that outcome, written before the run, was that the headline falls; we apply it. What survives is four preconditions, each measured on the same 63 cells, that a stratified evaluation must satisfy before its conditional number is about the regime: an unconditional advantage must exist to be inflated (absent in 13 cells); the labelling must be specified, since inside a frozen family of 75 labellings the estimand's reachable range crosses zero at every horizon on both networks; the two road networks must be reported separately, since they disagree wherever a threshold is involved; and the interval machinery must be calibrated, which here it is not — under a true null the nominal 95% interval excludes zero at 0.200, 0.290 and 0.445, that is 4.0, 5.8 and 8.9 times nominal. We report the null as a result rather than as a failure, and none of our conclusions rests on nominal interval exclusion.