On the Provable Emergence of Hierarchical Concept Structure in CLIP Embeddings
Abstract
Contrastive vision-language models such as CLIP are trained to align images and text, not to represent taxonomies. Yet empirical studies show that CLIP embeddings exhibit a hierarchical geometry aligned with the semantic hierarchy of concepts. A theoretical characterization of why such hierarchy emerges from a non-hierarchical objective remains open. In this paper, we formalize this phenomenon as Hierarchical Angular Separation (HAS), a geometric property whereby each non-leaf concept embedding is more similar to its descendants than to off-branch concepts. Empirically, we show that HAS already holds weakly in randomly initialized CLIP image embeddings, suggesting that hierarchical structure may be induced by the raw image geometry, while in text embeddings it emerges only after contrastive training. To explain this asymmetry between text and image embeddings, we analyze an unconstrained features model in which the image embeddings are fixed and only the text embeddings are trained. We show that when the image embeddings satisfy HAS with a positive margin, the global minimizer of the contrastive objective induces HAS in the text embeddings, thereby transferring hierarchical structure across modalities. The key mechanism is the hierarchical co-occurrence structure of image-text pairs, which biases text embeddings toward root-centered image prototypes. For a linear contrastive loss, we derive a closed-form solution that makes this mechanism explicit. We extend the analysis to the InfoNCE loss and establish analogous guarantees under sufficient regularization. Synthetic experiments validate the theory and show that HAS emerges even beyond the regimes covered by our analysis.