Approximation in Contrastive Representation Learning
Yuanfan Li ⋅ Zihan Zhang ⋅ Yiming Ying ⋅ Ding-Xuan Zhou
Abstract
Contrastive representation learning (CRL) is a fundamental paradigm for learning transferable representations. Recent work identifies its population target as the pointwise mutual information (PMI) $ \log\frac{p(\mathbf{x},\mathbf{y})}{p_\mathcal{X}(\mathbf{x})p_\mathcal{Y}(\mathbf{y})} $ between two modalities $\mathcal{X},\mathcal{Y}$, where $p$ is the joint density function and $p_\mathcal{X},p_\mathcal{Y}$ are the marginals. However, it remains unclear why standard inner-product scores $f(\mathbf{x})^\top g(\mathbf{y})$ can effectively approximate a generally nonseparable population target PMI. In this paper, we address this question from the aspect of approximation for CRL. By linking PMI to a compact operator, we show that $D$-dimensional inner-product models perform rank-$D$ spectral approximation, with error controlled by the spectral tail. We then study neural network realizations: in general settings, we show that deep neural network scoring functions are universally consistent, and the approximation error rates depend on the smoothness of the PMI function; under structured low-rank data models we obtain improved rates attained by moderate-width encoders. Finally, we establish a fast calibration rate from contrastive excess risk to downstream retrieval error under a density-ratio Tsybakov noise condition.
Chat is not available.
Successful Page Load