Emergent Visual Thinking in Text-Only Reasoning through Multimodal Training
Abstract
Does multimodal training make a language model think visually, even when the input is text only? We test this with targeted causal ablations across 10 multimodal configurations from 6 model families, with matched text-only backbones and cross-family transfer controls. We contrast spatial-imagery tasks that invite mental scene construction with symbol-manipulation tasks solvable by symbolic rules alone. We identify a vision-language axis in hidden-state space and ablate its projection at inference. In Chameleon-7B, this ablation drops accuracy on BIG-Bench Hard (BBH) navigate by 27.6 percentage points while improving boolean expressions by 5.2 percentage points, a 32.8 percentage-point double dissociation. The effect is direction-specific and layer-localized. Together, these results reveal emergent visual thinking (EVT): text-only reasoning that causally depends on visual-spatial representations shaped by multimodal training. In dual-path unified models (BAGEL, Janus-Pro), the causal handle localizes to the generation pathway as a compact one-dimensional carrier. The understanding pathway also encodes spatial information in a higher-rank, probe-decodable form, but is not the causal handle. To separate decodability from causal use, we analyze concentration, alignment, and semantic location. These diagnostics explain why behavior, probing, and intervention can diverge. Multimodal training therefore reshapes the geometry of text-only reasoning, beyond changing how visual input is processed.