A Unified Perturbation Framework for Analyzing Leaderboard Stability and Manipulation
Abstract
Evaluation leaderboards such as LMArena play a central role in benchmarking large language models by aggregating pairwise human preferences into model rankings, but the robustness of these rankings is still poorly understood. We present a unified perturbation framework for analyzing Bradley–Terry leaderboards under structured data modifications using influence-based approximations. Our framework studies three match-level perturbations—dropping matches, adding matches, and flipping match outcomes—along with player removal. We evaluate their effects on top-k membership, global ranking consistency measured by Kendall’s tau, and confidence-interval uncertainty. Across Chatbot Arena and six additional pairwise-comparison datasets, we show that modern leaderboards are non-robust across all three objectives: targeted perturbations affecting less than 1% of the data can change the top-ranked model, degrade global ranking consistency, and alter confidence intervals. We summarize these effects using normalized dataset-level robustness scores that compare fragility across leaderboard designs. We further show that influence scores enable efficient targeted manipulation, promoting or demoting specific models with fewer actions than prior manipulation baselines, while also identifying additional matchups that reduce uncertainty for target models. Finally, player-removal analysis shows that removing influential models can induce broad reordering, highlighting model deprecation as a source of the leaderboard illusion. These findings reveal fundamental limitations of current leaderboard designs and motivate more robust evaluation protocols.