How Far Is Too Far? Object Recognition Declines Monotonically with Semantic Distance
Abstract
Vision models are sensitive to object-background context, yet it remains unclear how recognition changes as the semantic distance between objects and backgrounds increases. We introduce ImageNet-OOC1k, a benchmark enabling continuous control of object-background relationships via a consensus semantic dissimilarity axis. This formulation reveals a highly consistent monotonic decline in recognition accuracy across 35 models spanning convolutional neural networks, Vision Transformers, and vision-language models. Internal analyses of Vision Transformers reveal aggregation failure, where object evidence is present at the patch level but not consolidated into the final prediction. Leveraging this characterization, a simple spatial re-ranking strategy recovers a subset of these errors and improves accuracy by up to 5.4 percentage points. These results establish context sensitivity as a continuous and predictable property of modern vision systems, with Vision Transformer analyses showing that failures can arise from limitations in aggregating competing signals rather than from missing object representations.