What Do Paired Vision–Language Benchmarks Measure? Cohort Cues, Metric Structure, and a Medical Case Study
Abstract
Vision language benchmarks are trusted to certify what models understand, and when a benchmark is suspected of rewarding one-sided shortcuts, the standard remedy is a paired design: score two images against two statements so that anything one-sided cancels. We show that passing a paired evaluation does not guarantee shortcut free measurement. The cancellation removes only additive effects, while cohort identity (which group each image came from, the very property that makes the paired answers differ) enters the score in the same functional form as the ability being measured, so the design requires the confound it cannot cancel. We derive two inexpensive diagnostics for this gap and validate them on published benchmarks: on BiVLC, one real-versus-generated direction carries a fifth of the image score (0.528 to 0.422) while the paired interaction barely moves (0.926 to 0.923); on Winoground, which admits no cohort axis, the same tests correctly find nothing. In addition, a controlled medical VQA case study traces the consequences for a concrete safety task by detecting when a question arrives with the wrong image: generator confidence is at chance (AUROC 0.502) and sampling-based uncertainty barely improves on it, one frozen BiomedCLIP alignment feature lowers unsafe answering from 0.542 to 0.376 (Δ = −0.166; 95% CI [−0.270, −0.077]), yet detection decays from 0.701 to 0.489 as the substituted image approaches the source, with a failure region near similarity 0.87, verified by two independent encoders. Therefore, these findings suggest that pairing only removes one class of shortcut, and that reporting these diagnostics alongside paired benchmarks would make explicit what such scores do and do not certify.