What Entropy-Based View Selection Cannot See: An Audit of Test-Time Adaptation for Vision-Language Models
Abstract
Vision-language models such as CLIP achieve strong zero-shot classification performance, but their accuracy can degrade under distribution shifts encountered in real-world deployment. A widely used remedy creates many randomly cropped and flipped copies of the single test image, called views, and combines their predictions. Because some views miss the object entirely, such a method must decide which views to trust. Methods in this family answer the same way: they measure how concentrated each view’s predicted class distribution is, a quantity called prediction entropy, and keep the views the model is most certain about. Subsequent research has refined this scoring function repeatedly, while leaving unexamined the information the score is computed from, namely the model’s own output. We examine that choice. Holding each stage of the pipeline fixed in turn, we first show that selection rather than the gradient update produces almost all of the accuracy these methods gain. We then ask whether the entropy score can identify a useful view in the cases where a better choice would help: images the method gets wrong although at least one view already predicts the correct class. On exactly those cases the score is no better than picking views at random. This holds for two different image encoders, under a control that rules out the obvious statistical artefact. An exact analysis of the entropy objective explains why, and a matched comparison isolates the missing ingredient: the score reports how certain a view is, never which class it favours. The pipeline nevertheless holds an unused signal. The augmentation procedure sampled each crop, so it already knows how much of the image every view contains, without any label. Ranking views by that quantity improves accuracy on 14 of 15 benchmark datasets, and a measurement taken before the method runs predicts which datasets benefit.