Same Models, Different Winners: How Evaluation Design Changes Model Selection in Financial Volatility Forecasting
Abstract
It is standard practice in financial forecasting to rank models by their loss on a single evaluation: one market, one error metric, one temporal split, one rule for averaging across assets. Few studies check whether that ranking survives an equally defensible choice of any of those four, and fewer still ask whether the leading model is separable from the rest at all. We generate every forecast panel once for six volatility models on 150 series drawn from crypto futures, Indian equities and US-listed equities and ETFs, freeze those panels, and score them under 1,044 evaluation protocols. Within a single market, changing only the volatility regime, the temporal backtest design, the loss function or the data-quality filter changes which model ranks first in roughly half of matched comparisons. Two of these alter no forecast at all, and they alone change the leader about twice as often as seed-to-seed variation does, even though such changes correspond to legitimate shifts in the evaluation estimand rather than to noise. The models themselves sit close together: the model confidence set retains most of the suite in every market, so the evaluation is resolving near-ties among models the evidence does not separate. We suggest that a leaderboard in this regime be presented together with the set of models the evidence cannot separate and with the ranking’s stability over a stated space of protocols .