Before Interpreting an AUROC: Four Silent Measurement Failures in Zero-Shot Clinical LLM Evaluation
Arianna Francesconi
Abstract
We evaluate six open language models for zero-shot prediction of 30-day hospital readmission from MIMIC-IV discharge summaries. Our initial analysis appeared to support a simple negative result: performance was near chance and several models collapsed to one class. An audit instead identified four independent failures in the evaluation pipeline. First-token verbalizer scores were read before some models could emit an answer; median answer-token mass was as low as $10^{-18}$. Generate-and-parse did not provide a general repair: 67.5\% of one model's outputs contained no yes/no token, while a catch-all rule determined 27.8\% of another's labels. A supervised baseline was joined to the wrong cohort by row order, and unconstrained Platt scaling reversed rankings for three models. After interface-aware rescoring, GPT-OSS-20B and GPT-OSS-120B reached AUROCs of 0.609 and 0.599, respectively, on a balanced benchmark, contradicting the original collapse claim but not establishing clinical utility. The central negative result is thus methodological: plausible metrics can survive multiple silent violations of their measurement assumptions. We provide diagnostics, boundary conditions, and a minimal reporting checklist intended to make these failures observable.
Chat is not available.
Successful Page Load