Forced Orders: What LLM Leaderboards Hide About Model Comparisons
Zonglin Di ⋅ Berk Ustun ⋅ Yang Liu
Abstract
LLM benchmarks are usually consumed as scalar leaderboards, where adjacent ranks are read as adjacent capability differences. We argue that this display should be complemented rather than simply trusted or discarded: benchmark scores are useful summaries, but a sorted score table can overstate which model pairs are actually distinguishable. This limitation is especially clear for accuracy- and success-rate leaderboards, whose averages hide task-level structure, and it also appears in pairwise systems such as Elo-style ratings when pairwise evidence is compressed into a total order. We propose LLM tier ranking as a task-level reporting framework for benchmarks. Given benchmark outputs, we construct pairwise wins, losses, and ties between models on each task and aggregate them into ordered tiers using Selective Preference Aggregation (SPA). Cross-tier comparisons are reported only when supported by sufficient task-level agreement; models in the same tier are left unresolved, not declared equal. Across AIME 2025, $\tau$-bench airline, xbench DeepSearch, GDPVal, SWE-Bench Verified, and Chatbot Arena, tier ranking often preserves useful broad capability classes while abstaining from unsupported fine-grained distinctions among frontier models. In Chatbot Arena perturbation studies, tier outputs are more stable than Elo-style total orders, expressing weaker evidence through later tier separation rather than rank reshuffling. We recommend reporting tiered leaderboards, solution paths, top-tier size, singleton-winner thresholds, and pairwise resolution statistics alongside scalar scores.
Chat is not available.
Successful Page Load