Benchmark Error Tracks the Perturbation, Not the Construct
Hyunoh Yeo
Abstract
Benchmark statistics are routinely accepted as quantifying model quality, but a low error certifies predictive accuracy, not dependence on the construct the benchmark is taken to measure. We treat this as a measurement-validity problem and propose an evaluation profile $P_m=(B_m,R_m,S_m)$ that supplements benchmark error $B_m$ with two auditable quantities, the alignment $R_m$ between predictions and a declared score $z^{\ast}$, and the perturbation coupling $S_m$, which asks whether benchmark error responds when $z^{\ast}$ is displaced and holds when $z^{\ast}$ is preserved. We instantiate the profile where the residual displacement of the score can be measured and not assumed, in cosmological parameter inference from simulated maps, where the simulator supplies ground truth and a physically motivated space of perturbations on the map. Auditing public third-party pretrained CNNs, we find benchmark-construct decoupling. At matched image-space magnitude, a perturbation that displaces $z^{\ast}$ eight times as far does not produce higher observed error, and across a sweep of perturbation strengths degradation tracks the size of the input change (Pearson $r=0.89$) and not the score displacement ($r=0.43$). A score-preserving filter that moves $z^{\ast}$ by $0.04$ of its own spread still raises benchmark error by $1.58\times$, $[1.38,1.79]$. The prediction itself is nonetheless strongly associated with the score (partial $R^2=0.62$ for the best-predicted parameter), so the benchmark grades the model by a robustness the score does not govern. The pattern replicates on an independently trained model from a different simulation suite, and a controlled pair of models with matched benchmark error shows the same separation the benchmark cannot make. We argue that evaluation should report alignment and coupling alongside error.
Chat is not available.
Successful Page Load