Validation Without Warning: Agreement-on-the-Line for Forecasting and a Label-Free Test of When It Breaks
Shresth Jaiswal
Abstract
Agreement-on-the-line (AGL) estimates the out-of-distribution accuracy of a model population from unlabelled data. It is established for classification, and it is known to fail for foundation models evaluated zero-shot, without an account of why. We derive the forecasting analogue under squared error. In terms of the excess errors relative to the Bayes forecaster, pairwise disagreement satisfies $D_{ij}=\Delta_i+\Delta_j-2\rho_{ij}\sqrt{\Delta_i\Delta_j}$ exactly, so a uniform change of the excess-error correlation $\rho$ across a shift moves the intercept of the agreement line, and a change that covaries with model loss moves its slope. Double-centring the disagreement matrix cancels the losses and yields a label-free, oracle-invariant statistic of the between-model error structure. On a controlled oscillator with a Bayes oracle, 144 trained forecasters satisfy AGL under an in-support shift, whereas across a bifurcation the agreement slope is twice the accuracy slope and $\rho$ rises from $0.08$ to $0.60$ with the rise concentrated on the better models. A pre-registered control over seven non-bifurcating shifts shows that the statistic tracks failure on trained models and tracks the inputs on 16 zero-shot foundation models (Chronos, TimesFM, Moirai): it responds to shifts that leave the failure structure unchanged and is silent on the non-noise shift that alters it. On real electrocardiograms a trained population reproduces the accuracy line under no shift, whereas under every real shift, a change of patient included, the accuracy line collapses while the agreement line persists. Agreement between zero-shot forecasters is therefore evidence about their inputs rather than about their errors, and label-free evaluation built on agreement requires an input-shift control before use.
Chat is not available.
Successful Page Load