Alignment Is Not Enough for Safe Medical LLM Evaluation
Abstract
LLMs-as-judges (LLJs) are increasingly used to reduce the cost of expert review in medical evaluation. However, high alignment with clinicians does not by itself establish validity or safety. This position paper argues that current reporting practices for medical LLJs are insufficient: LLJ-clinician alignment should not be reported as standalone evidence that a judge is reliable, clinically grounded, or safe for downstream use. We identify four risks: unstable evaluations across prompts, runs, and judge models; opaque scores that hide clinically meaningful errors; weak validity under clinical uncertainty and limited clinician agreement; and downstream safety or data contamination risks when flawed scores influence model selection or synthetic data filtering. By re-examining LLJ evaluation setups used in prior medical studies, we show that judges treated as practical evaluators in the literature can still reject clinically equivalent diagnoses, overrate incomplete or unsafe answers, reward unsupported clinical notes, and penalize appropriate treatment plans. We propose a protocol for medical LLJ evaluation that requires task-specific codebooks, expert-annotated validation sets, codebook-derived prompts, robustness testing, clinical error analysis, uncertainty reporting, and transparent disclosure before LLJ scores are used for medical benchmarking or decision support.