When Evaluation Goes Blind: Auditing Patch-Grid Failure Modes in Vision Transformers
Abstract
Automated vision-language evaluators are measurement instruments: their scores support model selection, dataset filtering, and benchmark claims. Yet their architectural blind spots are rarely considered when assessing evaluation validity. We present a black-box stress test that adds equal-energy checker patterns while varying cell width, resize ratio, injection stage, and spatial alignment. Across seven ViT-based evaluators, we find sharp drops in sensitivity at specific checker sizes determined by the model’s patch grid. A simple patch-grid rule predicts how these sensitivity drops shift under changes in resizing and model configuration across five encoders, and a separate SigLIP test follows the same prediction. Shifting the checker pattern relative to the patch grid reveals that sensitivity drops for exact- and super-cell patterns largely disappear when checker and patch boundaries no longer align, whereas the sub-even effect persists. The effect also propagates through one public aesthetic-scoring model, while two stride-32 convolutional controls with RN50 and ConvNeXt-B show much weaker effects. These findings reveal a reproducible evaluation failure mode: visually structured perturbations with the same energy can produce sharply different evaluator responses depending only on how they align with the model’s patch grid.