Spatial Grounding or Clinical Shortcut? Auditing Multimodal Evidence for Glaucoma Visual-Field Prediction
Abstract
Multimodal medical models are often credited with exploiting complementary clinical evidence when additional modalities improve predictive performance. However, such gains may also arise because the primary visual representation fails to encode information already present in the image. We study this distinction in glaucoma pointwise visual-field prediction from fundus photographs with optional OCT, demographics, intraocular pressure, and clinical text. We introduce a representation-conditioned audit that combines controlled spatial grounding, matched modality comparisons, leakage controls, and patient-disjoint validation. The analysis shows that learned spatial retinal representations substantially improve visual-field prediction and recover OCT-related structural information, while the incremental value of the six OCT summary metrics becomes statistically indistinguishable from zero under the end-to-end spatial representation in the evaluated cohort. Clinical-text gains are likewise not statistically robust after redaction and multi-seed controls. These findings motivate conditional evidence acquisition, reducing OCT use by 14–30% without a statistically detectable difference from always-acquire validation performance. More broadly, multimodal gains should be interpreted relative to what the underlying representation already captures, rather than attributed automatically to additional modalities.