Linguistic and Orientation Anchoring: Diagnosing VLM Failures in Fine-Grained Visual Matching
Abstract
Vision-language models (VLMs) excel across many multimodal tasks but still struggle with fine-grained visual comparison. Prior work attributes this weakness to linguistic anchoring: models map visual tokens to interpretable linguistic representations and reason in the linguistic space, potentially overlooking details that are difficult to verbalize. We first analyze frontier proprietary VLMs on paired matching. We select these latest and most capable models to determine whether the problem persists at the current capability frontier. Even without useful semantic labels, they continue to verbalize local structures, and random rotations and scaling cause substantial accuracy drops, revealing a failure mode we call orientation anchoring. Because proprietary models do not expose their weights or intermediate states, however, they cannot support mechanistic interpretability analysis. We therefore turn to the open-source Qwen and Gemma models for controlled fine-tuning and internal analysis. Fine-tuning on only 1,000 synthetic rotated pairs largely mitigates orientation anchoring and generalizes to unseen complexity levels and shape families, indicating that it is a learned shortcut. Visual-representation probing and logit-lens analysis further reveal two distinct adaptations: fine-tuned Qwen preserves fine-grained information from the original visual representation while remaining linguistically interpretable, whereas Gemma achieves comparable performance without preserving the original visual correspondence and with weaker linguistic separability.