Ranking Under Biased Missingness: When Model Rank Is Not Identifiable from Sparse Leaderboards
Jung Min Kang
Abstract
A leaderboard rank implicitly claims that the reported scores order two models; on sparse leaderboards, that claim is usually taken for granted rather than checked. Public robotics and embodied-AI leaderboards report scores for only subsets of benchmarks. We study a cross-benchmark matrix spanning LIBERO, CALVIN, Meta-World, and LIBERO-Plus with 162 deduplicated models, 18 benchmark items, and 781 observed scores (26.8% density). Reporting coverage is negatively correlated with fitted item difficulty ($r=-0.639$), but this association does not identify the reporting mechanism. We fit a penalized two-parameter logistic item-response model and bootstrap observed cells ($B=1000$). The resulting intervals quantify pipeline sensitivity, not coverage-guaranteed uncertainty: 36.4% of model pairs are unresolved by the pairwise criterion on the unmasked matrix. The maximum divergence between latent-trait and average-based ranks is $\Delta=77$; the corresponding model's rank interval spans 58 positions. Under additional masking at 50% and 90%, we distinguish a fixed-denominator rate that includes unavailable pairs from a rate conditional on surviving models. At 90% masking these rates are 91.6% and 45.6%, respectively. Rank self-consistency with each estimator's own unmasked ranking also changes ordering: $\Delta\rho=-0.0031$ at 50% masking and $\Delta\rho=+0.106$ at 90%. These are outcomes for the sampled masks, not evidence of a stable regime boundary or superior ranking accuracy. Before a sparse ranking is trusted, it should report coverage and state which comparisons it can make, which it cannot resolve, and which it cannot make at all.
Chat is not available.
Successful Page Load