Patched, Not Repaired: A Low Benign False Positive Rate Certifies the Region, Not the Model
Abstract
A low benign false-positive rate establishes that the evaluated manifestations of a failure have been suppressed; it does not establish that the failure mode has been removed. The same segmentation model, deployed in an East African motor insurance claims pipeline, yields a benign false-positive rate of 0.12% on one set of images containing no broken glass and 97.5% on another, and neither rate is wrong: what differs is the region of input space each set samples. The response is not semantic, since the model fires on tree canopy off the vehicle and on synthetic noise containing no object at all. Hard negative correction then repairs four families of a 280 image stress benchmark, halves a fifth, and leaves fur untouched. Supplying fur drives held-out fur below detection in all three seeds while the unsupplied synthetic probe is unmoved in every epoch of every run; we do not claim that survivor is the same response, only that no evaluation we ran can tell. The intervention is also not selective: on 52 dry grass photographs, a post hoc family never supplied as a negative, firing falls from 44% to 8% while the matched control rises to 60%, in the same direction in four of four runs. Thirty images retuned the model's operating point rather than teaching it a distinction, and no conventional evaluation distinguishes the two. We set out what we would report instead, and will release the probes, the evaluation code and the benchmark specification.