Beyond Pairwise Supervision: Spectral Characteristic Matching for Data-Efficient Multimodal Alignment
Xinchao Wang ⋅ Shuang Li ⋅ Wei Chen ⋅ Jingxuan Kang ⋅ Fuzhen Zhuang ⋅ deqing wang
Abstract
Multimodal alignment under scarce paired supervision aims to align independently pretrained unimodal encoders using limited image--text pairs, offering a practical alternative when large-scale joint pretraining is infeasible. Such scenarios call for lightweight objectives that go beyond local correspondences and capture distributional mismatch between heterogeneous embeddings. Existing methods often derive supervision from limited pairs through contrastive objectives or geometry-preserving regularization, encouraging instance-level correspondence or structural consistency but providing limited direct supervision on global cross-modal mismatch. This motivates a distribution-level objective that complements paired supervision by comparing image and text embedding distributions. In this paper, we propose $\textbf{Spec}$tral Characteristic Distribution $\textbf{Align}$ment ($\textbf{SpecAlign}$), a lightweight spectral alignment framework that reduces cross-modal mismatch by comparing image and text characteristic functions over learnable spectral queries. Since characteristic functions uniquely determine probability distributions, SpecAlign provides a principled signal for capturing multi-scale discrepancies beyond pointwise alignment. To obtain informative and stable queries, SpecAlign employs a Structured Direction--Radius Sampler (SDRS) that decouples direction and radius modeling. Extensive evaluations on transfer, retrieval, robustness, and representation analyses demonstrate consistent gains, validating spectral distribution-level supervision for data-efficient multimodal alignment.
Chat is not available.
Successful Page Load