A Unified Spectral Theory of Multimodal Losses
Abstract
Modern vision--language models are trained with very different objectives: CLIP-style contrastive alignment, SigLIP-style match/no-match prediction, caption-style token generation (Flamingo, LLaVA), and latent-space prediction (VL-JEPA). Each has been claimed as best-in-class for its target task. Yet practitioners lack clear guidelines for choosing between them, and existing identifiability theory analyzes only the contrastive case. In this work, we unveil the core statistical mechanisms that distinguish each objective family using a unified spectral analysis. By leveraging closed-form solutions for the linearized losses under a latent structure in which modality-specific nuisances may spuriously carry shared semantic content, we characterize what each objective recovers from the data. We demonstrate that contrastive alignment, match/no-match prediction, and latent-space prediction collapse to the same cross-modal solution, while directed token-space generation recovers a different one. Our findings indicate that in scenarios where one modality is heavily contaminated by spurious nuisance, token-space generation is preferable because it imposes a strictly weaker, one-sided alignment condition. In scenarios where two modalities are both clean, retrieval is maximized by the latent-prediction objective whose target matches the query modality. These results clarify the trade-offs between the four objective families. We validate our theoretical predictions on numerical, controlled synthetic, and real (Flickr30k, COCO, CC3M) data.