Beyond Eigenfunctions: Divergence Principal Functions for Representation Learning
Ritabrata Ray ⋅ Sahil Dharod ⋅ Burak Varıcı ⋅ Nicholas Boffi ⋅ Pradeep Ravikumar
Abstract
Recent work has shown that contrastive representation learning can be understood as estimating a positive-pair, PMI-like kernel, and that spectral factorization of this kernel yields useful eigenfunction representations. This paper asks what happens beyond squared-error spectral geometry. Modern representation learning objectives are rarely pure $\ell_2$ kernel-approximation objectives: they use conditional KL, InfoNCE, JS/NCE, Brier, Hellinger, least-squares ratio fitting, and other statistical discrepancies. We show that these discrepancies do not merely provide alternative estimators of the same affinity; they induce different local geometries for the extraction of low-rank representations. To formalize this, we introduce divergence principal functions, a geometry-aware generalization of eigenfunctions. Eigenfunctions are recovered as the special case corresponding to squared-error geometry. For smooth discrepancies, we show that divergence principal functions locally solve a weighted spectral approximation problem, with weights given by the curvature of the divergence at the target affinity. This provides a loss-consistent bridge from positive-pair learning objectives to representation geometry and clarifies when spectral/eigenfunction representations are appropriate and when they are mismatched to the training objective.
Chat is not available.
Successful Page Load