Aligning Language Model Benchmarks with Pairwise Preferences
Abstract
Language model benchmarks are pervasive and computationally-efficient proxies for real-world downstream performance. However, many recent works find that benchmarks often fail to predict downstream utility. While some works have begun diagnosing sources of misalignment, there remain no ways to systematically update benchmarks to align their scores with downstream usage. Towards bridging this gap, we introduce and study \textit{benchmark alignment}, where we use information about downstream model performance to automatically update benchmarks, aiming to produce new static benchmarks that predict model pairwise rankings on downstream tasks. Our experiments involving 4576 language models and 6 benchmarks show that benchmark items can successfully be reweighted to predict downstream performance for unseen models, even generalizing across model scales in most cases. And while naive alignment unsurprisingly requires large numbers of models and benchmark questions, we find that proper use of downstream data can often also enable alignment, even using responses from as few as 20 models. Overall, our work takes a step towards efficiently aligning benchmark development with downstream tasks.