From Topology to Retrieval: Decoding Embedding Spaces with Unified Signatures
Abstract
Studying the shape of text embedding spaces enhances our understanding of model behavior and how geometry affects downstream task performance. While a range of topological and geometric metrics have been explored, they are typically analyzed in isolation, and the connections between them remain largely unknown. In this paper, we present Unified Topological Signatures (UTS), a holistic framework for characterizing the structure of embedding spaces. We evaluate our method across a wide range of text embedding models and retrieval datasets, finding that UTS outperform individual metrics in downstream performance prediction. Through several experiments, we demonstrate that UTS are superior at characterizing embedding spaces, can predict model-specific properties and bias from geometry, and reveal novel insights into representational similarity that challenge the Platonic Representation Hypothesis.