Retrieval Rankings Do Not Directly Identify Multimodal Latent Space Alignment
Eleonora Grassucci ⋅ Danilo Comminiello
Abstract
Multimodal representation models are often called aligned when paired items rank highly in cross-modal retrieval. Nevertheless, a rank records relative order in a candidate set, while claims about a shared latent space concern its geometry, orientation, and global organization. We ask whether retrieval rankings contain enough information to support those broader claims. We introduce two rank-equivalent transformations and evaluate the geometry of the resulting spaces. Across four CLIP backbones and two retrieval benchmarks, the ranking remains fixed while the distribution of embeddings in the latent space varies consistently according to several geometry-aware metrics.
Chat is not available.
Successful Page Load