Permute-then-Adapt: Weak-to-Strong Contrastive Image--Text Adaptation
Jinhao Li ⋅ Sarah Erfani ⋅ Lei Feng ⋅ Guangrui Li ⋅ James Bailey ⋅ Feng Liu
Abstract
Image--text alignment models such as CLIP are typically trained with contrastive learning, where unpaired examples are treated as strict negatives. While this assumption introduces noise by penalising semantically related pairs, existing solutions often rely on heuristic similarity thresholds that lack statistical grounding and sensitivity to dataset-specific noise distributions. In this paper, we propose **Permute-then-Adapt (PTA)**, a framework that leverages efficient *weak models* to identify missed positives, while enforcing *statistical control* to reject random associations. Unlike standard distillation or arbitrary thresholding, PTA estimates the null distribution of the weak teacher's similarities via permutation testing, ensuring selected pairs are *statistically distinguishable from the distribution of random pairings* at level $\alpha$. The calibrated positive set is then trained against by a single multi-positive contrastive objective: each anchor maximises the total probability mass it assigns to its positive set. We further demonstrate that the calibration can be pre-computed offline using the weak model, so training overhead is negligible compared to standard baselines. Extensive experiments show that PTA consistently outperforms heuristic soft-label approaches on object recognition and cross-modal retrieval, while exhibiting superior data efficiency.
Chat is not available.
Successful Page Load