Same Eye, Different Scanner: A Decision-Reliability Benchmark for Frontier VLMs and OCT Foundation Models in Glaucoma
Hao Li ⋅ Jared Joselowitz ⋅ Zachary Ellis ⋅ Yuxiang Zhou ⋅ Samuel Thio ⋅ Aisling Higham ⋅ Yajie He
Abstract
A glaucoma-OCT reader used across providers should give the same eye the same diagnosis whatever scanner acquired the image. In practice, however, OCT measurements are not interchangeable across scanners, so a model's decision can change even though the eye is unchanged. We ask whether prompted frontier vision-language models (VLMs) and specialist OCT foundation models (FMs) keep a faithful, calibrated decision when the scanner changes. We build a benchmark that applies the scanner change as a same-eye image transform calibrated to measured cross-scanner disagreement. Because each eye is scored before and after, the scanner effect is de-confounded from patient differences while the images and labels stay real. We score the decision at a fixed operating point (a single decision threshold) by its decision-flip rate, change in sensitivity, and expected calibration error (ECE), with bootstrap confidence intervals (CIs) on the decision metrics. Four results follow. (1) Prompted frontier VLMs are the weakest and most overconfident above-chance readers on the raw B-scan (AUC 0.61–0.69). Their ECE is roughly $5$–$10\times$ the best FM's for the Gemini and GPT families (Claude Opus 5 a better-calibrated exception); smaller open-weights and medical-specialist VLMs are near chance (0.41–0.59). (2) Rendered as a color thickness map, the stronger VLMs reach the level of the strong generic-FM and CNN readers (AUC 0.83–0.84), far above their raw-B-scan scores. The small open models do not recover, so what lets a VLM read the scan is the display format together with a capable base, not model size alone (a cross-dataset comparison). (3) A scanner change can break the decision while discrimination is intact: $\Delta\mathrm{AUC}\approx 0$ while the diagnostic decision flips and specificity can collapse, invisible to AUC-only evaluation. (4) Reliability tracks the input representation and pretraining breadth, not model size: across a 26M–197M CNN sweep the largest model is steadiest on neither axis, and on the thickness-map axis a 98k-parameter from-scratch CNN matches 300M-parameter models on decision stability.
Chat is not available.
Successful Page Load