ITO: Multi-View Alignment and Training-Time Fusion for Image-Text Pretraining
Abstract
Image--text contrastive pretraining has become a dominant paradigm for visual representation learning, yet existing methods often yield representations that remain partially organized by modality rather than by semantics. We propose \textbf{ITO}, a framework addressing this limitation through two complementary mechanisms with distinct roles. \emph{Multimodal multiple alignment} enriches supervision by constructing diverse cross-modal correspondences from multi-view image augmentations, providing the primary source of discriminative gain. A lightweight \emph{training-time multimodal fusion} module then acts as a geometric regularizer, encouraging the encoders to produce features that are compatible under fusion and thereby reducing modality-induced separation. Crucially, the fusion module is discarded at inference, preserving the efficiency of standard dual-encoder architectures while incurring training-time overhead only. Extensive experiments across pretraining scales from millions to billions of image--text pairs show that ITO consistently outperforms strong baselines on classification, retrieval, and multimodal benchmarks. Our analysis further reveals that beyond accuracy gains, training-time fusion plays a stabilizing role in optimization, mitigating the late-stage overfitting commonly observed in aggressive contrastive learning.