Diagnosing and Repairing Visual Collapse in Compact Medical Multimodal LLMs
Abstract
Compact medical multimodal models are the most realistic path to clinical deployment, yet they often produce fluent answers that are not actually grounded in the image. We show that this failure has a clear internal signature. In five compact medical backbones and across six medical VQA benchmarks, visual tokens lose their spatial diversity shortly after entering the language backbone, their residual updates are an order of magnitude smaller than text residuals, and attention on them freezes on content free regions of the image. Zeroing the visual residual stream barely changes the output, while zeroing the text residual stream destroys it. We give a short analytical account that links the collapse of visual similarity to the residual dominance ratio, and we propose VGrip, a set of three lightweight losses that act on the three failure points without modifying the backbone architecture or the inference procedure. We also introduce the Visual Dependency Score, a simple and benchmark agnostic measure of how much a model actually uses the image. On Qwen2.5-VL-7B, VGrip lifts accuracy by 4.7 points on average across the six benchmarks and more than doubles the Visual Dependency Score over a strong LoRA baseline, with gains that are stable from two to eight billion parameters.