Open-Weight MLLMs Cannot Compare What They Cannot Name
Abstract
Multimodal Large Language Models (MLLMs) are increasingly capable of understanding images and performing complex visual reasoning, suggesting the possibility of using them to automate biological image analysis. This need is becoming increasingly urgent as advances in technology continue to generate ever larger and more diverse microscopy datasets, already exceeding the scale at which manual visual analysis is feasible. With the final aim to use MLLMs for visual comparison of complex phenotypes, we investigate a very simple baseline: comparison of the size of body parts in synthetic images of ants. We find that state-of-the-art open-weight MLLMs fail for some body parts, but not for others, even though the comparison itself is visually trivial for all of them. We evaluate different in-context learning strategies and analyze performance on the related object recognition task to come to the hypothesis that models struggle to visually ground semantic concepts they cannot recognize in the image. Further analysis of the vision encoder suggests that fine-grained visual representations are insufficiently preserved in cross-modal training unless supported by semantic associations. Our results establish a surprising limitation of current MLLMs: semantic recognition can gate simple visual comparisons that do not intrinsically require semantic understanding.