Detection Without Diagnosis in Small Evaluators
Abstract
Excluding a near-reject-all Phi evaluator, a field checklist detects 47.0% of controlled faults at 95.8% specificity. A generic score detects 18.9% at 90.3%. In the full four-model pool, the checklist detects 458 of 768 faults but names the changed field in only 185 detections. Phi raises apparent recall by failing unrelated fields and passes only 2.1% of controls. We test the same diagnosis gap with GPT-5.6 Sol, where every detection names the changed field in three complete checklist runs. Mean target recall is 89.1%, with a range from 88.5% to 89.6%, and all 144 run-control judgments pass. Pairwise field-vector disagreement ranges from 1.7% to 3.3%. Supportive acknowledgment behaves as a construct boundary rather than a demonstrated evaluator blind spot. Its omission is accepted often under every interface, so that result depends on treating acknowledgment as a required field. These results come from 240 fictional, non-graphic help-seeking replies judged through the same prompts and parsers.