The Screen Decides the Verdict: Operating Characteristics of Model-Exclusion Rules in LLM Evaluation
Abstract
An evaluation of hosted language models can screen models for run-to-run stability and then vote among the survivors. We audit a preregistered evaluation in which this two-step pipeline determined the result. Four models answered a yes/no concern question about 20 behaviours written to be ambiguous, each attributed to a neutral actor or to one of six personas. The registered stability gate required a Pearson correlation above .80 between two runs over 14 conditions (two scenarios under seven frames) with 20 calls each. It excluded the model whose rates moved least on the rate scale (mean absolute change .018) and retained the model whose rates moved most (.129). A simulated model with the excluded model’s rates and no drift failed the gate in 83.9% of bootstrap draws. Across 252 simulated designs, the gate excluded models that had not drifted at a median rate of .763. At the study’s design, with an assumed spread between conditions of .5 logits, we recalibrated the threshold to a 5% false-exclusion rate. Power against an upward one-logit shift at a baseline rate of .50 (moving it to about .73) was .066. A registered hypothesis, that persona frames make a model’s answers less consistent within a condition, was to be supported if at least 3 of the 4 models showed this. After the exclusion this rule required 3 of 3, and the realised outcome, two passes and one failure, met neither the support rule nor the paired rule that would refute it. A non-confirmatory sensitivity analysis that restored the excluded model met the support rule. We study one protocol and make no prevalence claim. We show that fixed-count support and refutation rules that give every full-panel outcome exactly one verdict leave some outcomes with no verdict after a removal, and that restating the registered thresholds as fractions of survivors does not repair our case. Screens and vote rules should therefore be specified and simulated together.