Rethinking Projector Training in Multimodal LLMs
Fabian Gröger ⋅ Naël Ouerghemi ⋅ Shuo Wen ⋅ Maria Brbic
Abstract
Multimodal LLMs are typically trained by first aligning a vision encoder with a language model via projector training on large image-caption datasets, followed by joint training of the full multimodal model. Despite its widespread use, the role of the projector training stage remains poorly understood. Does it learn fine-grained visual-language correspondences, or mainly maps visual features into a compatible input space for the language model? We find that (i) using only $10$ to $20\%$ of the data for training the projector already recovers most downstream performance gains, and (ii) projectors trained on different datasets can be linearly interpolated while largely preserving performance. These results suggest that projector training acts primarily as a coarse alignment within a large solution space. Motivated by these findings, we propose PORTAL, a training-free projector initialization that requires neither paired image-caption data nor gradient-based optimization, but only per-modality summary statistics. PORTAL computes an optimal-transport-based map between Gaussian approximations of the vision and language feature distributions on a shared principal subspace, enabling completely skipping projector pretraining. Despite using no paired data, while the standard approach relies on 0.5M image–caption pairs, PORTAL matches the default trained pipeline across six LLM backbones, two vision encoders, and 16 benchmarks, outperforming it in $3$ of $7$ (vision, LLM) settings, and remaining within $1.6$ average points across all other settings.
Chat is not available.
Successful Page Load