The Camera, the Split, and the Label: A Three-Check Audit of Glaucoma-Progression Prediction
Abstract
On the public GRAPE benchmark, deep networks report glaucoma-progression AUCs as high as 0.96. We audit the benchmark's published record and find that, at this sample size, none of the reported numbers can be distinguished from a trivial baseline. We re-evaluate all five published models under one fixed protocol of three checks: a device-only baseline (on GRAPE, the acquisition camera), a patient-grouped split, and a test that the input is disjoint from the window that defines the label. As a positive test, we predict the shipped label from the baseline visit, before its trend is computable. Image features from a frozen fundus model identify the camera at AUC 0.98-1.00, the camera alone predicts progression at 0.56-0.60, and a device-controlled fundus model does no better. Under a fair, device-controlled evaluation every reported number is no better than it, and one model's own released code falls from 0.87/0.96/0.92 to 0.43/0.54/0.55 (PLR2/PLR3/MD). The label is a trend fit to the visual-field series, and we introduce a circularity dial that measures how it is recovered from those defining statistics, with no image, as the input window approaches the window that defines it (AUC 0.90). The checks need only metadata GRAPE already includes: a device label, patient identifiers, the label's definition, and visit dates. We release a fixed split and a reporting template so the audit can be reused.