Measuring Severity
Abstract
Accuracy treats every error as equivalent. Two systems with identical accuracy are not truly equivalent if one makes plausible errors that a competent expert might also make, whereas the other generates outputs no human expert would ever produce. Current benchmarks fail to distinguish between these two failure modes. We argue that severity labels, expert judgments of the harm done by a particular wrong answer, constitute a missing evaluation primitive whose absence is felt most strongly in efficiency research. Efficiency methods such as compression, quantization, and adaptive inference are selected and compared on accuracy, yet what a compressed model forgets is not random: errors concentrate on particular instances and classes. Without severity labels, choosing the cheaper model is a blind trade between compute and harm. For this Grand Challenge, we argue severity is most urgent for efficient models, whose selection is a cost-benefit decision, and we aim to build evaluations that measure how often a model fails poorly in clinical prediction.