RTEB: An Overfitting-Resistant Benchmark for Embedding Model Evaluation
Sahil Verma ⋅ Minghan Li ⋅ Andrew Gaut ⋅ Yujie Qian ⋅ Kaidi Cao ⋅ Zhenmei Shi ⋅ Kenneth Enevoldsen ⋅ Roman Solomatin ⋅ Isaac Chung ⋅ Tom Aarsen ⋅ Apoorva Joshi ⋅ Emilia Garcia-Casademont ⋅ Michael Günther ⋅ Niklas Muennighoff ⋅ Yevhen Kostiuk ⋅ Fődi Zoltán ⋅ Tengyu Ma ⋅ Frank Liu
Abstract
Embedding models are central to modern information retrieval, retrieval-augmented generation, and downstream NLP systems, yet their evaluation remains unreliable: widely used benchmarks such as BEIR and MTEB rely on fully public datasets, making them vulnerable to overfitting that inflates scores without corresponding gains in real-world generalization. We introduce RTEB, a retrieval-focused embedding benchmark designed to address these limitations. Unlike MTEB's broad multi-task scope, RTEB is purpose-built for retrieval -- the setting most critical to RAG pipelines and agentic systems. RTEB emphasizes underrepresented yet high-impact domains, particularly code and agentic code retrieval, and keeps a large portion of its datasets private to structurally prevent direct test-set overfitting. We evaluate 20 models on 48 datasets (20 public, 28 private) spanning six domains and four languages. Rankings on RTEB diverge substantially from BEIR and MTEB ($\rho = 0.38$), agentic code retrieval emerges as the hardest domain with the best model reaching only 0.48 NDCG@10, and domain-specialized models fail to outperform strong generalists on any domain. Together, these design choices and findings position RTEB as a more reliable, discriminative, and future-proof benchmark for embedding model evaluation.
Chat is not available.
Successful Page Load