InfCLIP: Unified Data Valuation for CLIP Pretraining via Influence-Inspired Scoring
Abstract
Effective data valuation is essential for machine learning, especially for CLIP pretraining, where models are trained on massive and inherently noisy image-text datasets. Existing CLIP data selection methods often combine heuristic signals such as alignment, target similarity, and diversity, but lack a unified principle for estimating sample contribution. This can lead to inconsistent performance across data regimes and obscures how each sample contributes to downstream performance. In this work, we propose InfCLIP, a principled influence-based approach for data selection in CLIP pretraining. InfCLIP estimates the contribution of each training sample to a target task by leveraging feature representations and temperature-rescaled softmax distributions, enabling efficient computation without retraining or access to model internals. This formulation provides a unified view of alignment, informativeness, uniqueness, and target relevance within a single framework. Empirically, InfCLIP consistently outperforms existing data selection methods across a wide range of benchmarks, including zero-shot evaluation on ImageNet and distribution-shifted datasets. We also show that Self-InfCLIP enables efficient training-set analysis, including noise detection and memorization-aware data assessment, without additional retraining.