Multimodal LLMs Can Beat You at College Math, but They Can’t Count to 10: A Mechanistic Account
Neslihan Bulut ⋅ Amirhossein Farzam ⋅ Mohammadhossein Bateni ⋅ Vahab Mirrokni
Abstract
Multimodal large language models (MLLMs) struggle with elementary visual reasoning tasks such as counting. We study \emph{compositional counting} as a controlled testbed for localizing these failures. We find that queries marginalizing over an object attribute are substantially harder than conjunctive queries, even when the image, objects, and correct answer are identical. In contrast, models exceed $93\%$ accuracy on textual descriptions of the same scenes. To identify the mechanistic source of this failure, we formulate compositional counting as the interaction of scene representation, query-conditioned selection, and aggregation, and construct counterfactual interventions that isolate these operations. Our mechanistic analysis indicates that quantitative scene information is accessible in visual activations, but counting-ready representations emerge only in deeper layers. Patching these late representations into earlier layers significantly improves counting without retraining. These results identify a representation-timing bottleneck: MLLMs compute information useful for counting, but too late for query-conditioned reasoning to effectively use it. More broadly, our framework distinguishes whether task-relevant information is merely represented or available where downstream multimodal computation can use it.
Chat is not available.
Successful Page Load