When Do Cosine Prototypes Mislead? A Whitening-Aware Benchmark for Frozen-Feature Image Classification
Jian Ding ⋅ Su Yang
Abstract
Cosine nearest-class-mean (NCM) evaluation is a widely used default for closed-set frozen-feature image classification because it is deterministic, train-free, and easy to reproduce. This convenience hides a metric assumption: raw angular geometry alone suffices for benchmark conclusions. We show that the assumption changes model-selection outcomes, not just absolute accuracies: across nine datasets and six frozen backbones, cosine-only reporting changes which backbone ranks first on 2 of 9 datasets and incurs 6.71 percentage points of regret relative to a best-evaluator oracle; on fixed DINOv2-B features, the underlying cosine-to-LDA gap reaches 37.63 points on Aircraft. Motivated by the low-variance spectral directions where class-discriminative signal often appears in such high-gap cases, we call this evaluator dependence the Spectral Decision Gap (SDG). SDG quantifies the gap between raw cosine and train-validated whitening-aware reference evaluators, flagging when cosine-only reporting changes benchmark conclusions. A leave-one-dataset predictive protocol reaches Pearson $r=0.752$ for best-whitening disagreement and incurs only 0.56pp routing regret at a 5pp threshold. This paper provides a compact audit protocol and artifact suite for detecting, predicting, and reporting when cosine prototype conclusions are reliable in this setting. Code and artifacts are available anonymously at https://anonymous.4open.science/r/sdg-ed/.
Chat is not available.
Successful Page Load