Learning to See Through Language: An Exploration of Language Modeling's Effect on Visual Representations
Abstract
Does language modeling on image-text data truly improve vision, or does it merely adapt visual features for language alignment? Despite rapid progress in multimodal large language models (MLLMs), this question remains poorly understood. We present the first systematic analysis of visual representations in MLLMs by probing popular model families across a broad suite of dense visual tasks and comparing each MLLM to its exact pre-MLLM vision encoder. We find that language modeling largely improves visual representations for semantic tasks but can degrade performance on fine-grained geometric tasks, such as monocular depth. We further show that scaling the language model consistently improves the quality of visual representations. Given the improved MLLM representations, we next examine how to best extract and aggregate these features. Since no single layer provides a universally strong representation and concatenating features across layers is impractical, we propose a lightweight adapter that efficiently combines features across MLLM layers, turning frozen MLLMs into competitive visual backbones. Finally, we study whether MLLM representations can improve downstream applications, including text-to-image generation and inverse dynamics modeling, and show that they can outperform pre-MLLM visual features as general-purpose representations. Together, our results provide a representation-centric perspective for understanding and leveraging MLLMs as visual foundation models.