When Better Than Chance Is Not Enough: Net-Benefit Bands for LLM Verifiers
Abstract
A large language model (LLM) verifier is a second model that checks the output of a first model. It may separate correct from incorrect outputs yet still make a system worse if it rejects too many correct ones. We observed this problem in a biomedical metadata pipeline. We compared a baseline verifier with seven modifications. The rebuilt verifier improved Youden’s J from -10% to +11%. Here, J is the difference between how often the verifier flags incorrect and correct study-level calls; a positive value means that it flags errors more selectively. Filtering also raised the precision of the calls that remained from 34% to 39%. These improvements did not, by themselves, show that the verifier was useful for our evaluation objective. Decision-curve analysis weighs the benefit of removing false positives against the harm of removing true positives. The rebuilt verifier had the highest estimated net benefit among the tested strategies only for action thresholds between 29.5% and 39.2%. An action threshold expresses how an operator values these two types of error. Balanced accuracy, which gives equal weight to positive and negative cases, corresponded to a threshold of 14.4%, outside that range. In this setting, the alternatives were to leave the pipeline unchanged or to remove every positive call. The condition J>0 already tells us whether any useful range exists. The range still matters because its endpoints show how often the removed and retained calls are correct, and therefore which error trade-offs favor filtering. It was empty in 10.5% of 4,000 study-level bootstrap samples. We recommend reporting this deployment range, its uncertainty, and the intended operating point together with standard verifier scores.