Calibrating Score Banks for Earth-Observation Anomaly Detection: Multiplicity, Tail Shape, and Power
Abstract
Score banks are commonly used to detect heterogeneous Earth-observation anomalies, but combining them by their maximum raises the calibrated rejection threshold. We show that linear redundancy alone does not predict this aggregation cost. Using a frozen twelve-head bank built from matched image-only and multimodal EarthNet2021 forecast-residual features, we separate statistic selection, cube-level split-conformal calibration, and evaluation across 27 injected-anomaly scenarios. Although the bank spans only 1.82 effective linear directions, its maximum threshold exceeds the median individually calibrated head threshold by 0.98 variance-normalized units, compared with 0.27 predicted by effective rank. A joint-exceedance audit reveals a tail effective multiplicity of 3.32, while non-Gaussian upper tails account for the remaining 0.46 excess. This cost has practical consequences: cross-fitted selected heads outperform the maximum by 5.9-25.4 percentage points in seven scenarios, while matched-FPR comparisons show gains in six scenarios and losses in six. Max aggregation is therefore not uniformly beneficial, and its cost should be assessed using empirical joint-exceedance and tail-shape diagnostics rather than linear correlation alone.