Short-circuit Error Attribution Fails for Bottleneck Evaluation
Abstract
Evaluating which stage of a compound AI pipeline is responsible for its failures is central to targeted improvement. A widely used diagnostic, short-circuit error attribution, walks each failed case through the pipeline and blames the first stage whose output is wrong. We argue this fails to fulfill two natural motivations for bottleneck analysis: finding the stage that is intrinsically least reliable, and finding the stage whose correction would eliminate the most failures counterfactually. We formalize both notions and show, via a simple example and a general proposition, that short-circuit error attribution in general fails to perform either goal. As an alternative, we propose oracle substitution, a necessity test that substitutes in a corrected version of each stage's input and measures the resulting change in its own error rate, yielding an intrinsic error rate comparable across stages. We demonstrate the short-circuit error attribution and the proposed oracle substition using a simple four-stage Mini-ARC perception-reasoning pipeline (o4-mini throughout), where short-circuit error is shown to not accurately identify the bottlenecks.