Fidelity Ceilings Break Reliability Coefficients in Generative AI Evaluation
Mst Mousumi Rizia ⋅ Hanna Suominen
Abstract
We report a reader study in which the standard measure of reader reliability returned a verdict of failure on a study that had not failed. Four board-certified radiologists rated two medical inpainting models across 44 cases and four criteria; $73.1\%$ of judgements fell at the scale maximum. Median observed agreement was $0.616$ against a median chance-expected agreement of $0.631$, leaving Fleiss' $\kappa$ negative in six of eight strata, which reads as expert consensus worse than chance, and is not. We show this is a statement about the contrast rather than about the readers, using the same judgements and no model: recoding each rating pair as a within-case difference raises the agreement coefficient in four criteria of four, by a median of $+0.158$, and the increase survives holding out any single reader. The collapse survives every such refit too, so no single reader accounts for it. The mechanism is classical $\textemdash{}$ agreement coefficients estimate between-case discrimination, of which $3.9\%$ of the variance here remained $\textemdash{}$ but the regime is not: in a survey of ten recent reader studies of generative medical images, none reported a saturated instrument and seven used designs on which saturation is not measurable at all. We recommend reporting that lets a reader distinguish saturation from disagreement.
Chat is not available.
Successful Page Load