When Calibration Records Cannot Support Routing: A Denominator and Artifact Audit
Abstract
Low calibration error does not establish that verbalized confidence can support routing. We audit the evidence needed for that decision in two retained records. The first is nearly perfectly accurate and has low aggregate calibration error, but it contains only one observed error and its joint answer-confidence coverage varies sharply by endpoint. It can describe average confidence on this finite record, but cannot establish discrimination or selective risk. The second record omits item text, answers, gold labels, raw outputs, and verification evidence. Reproducible arithmetic on its supplied correctness field therefore does not become a verified behavioral estimate. Together, the cases change the operational decision from choosing a confidence threshold to redesigning the items, interface, and evidence record before any routing evaluation. We provide a compact machine-readable audit covering denominators, artifact status, calibration, discrimination, selective risk, uncertainty, and unidentified outcomes. The contribution is a bounded workflow for deciding what a calibration record can support, not a validated benchmark or a general claim about verbalized confidence.