RankShift: Evaluating Ranking Instability Across Benchmark Dimensions in Document Parsing
Abstract
Document parsing benchmarks are often treated as stable measures of model quality, yet rankings may depend substantially on benchmark design choices. In this work, we present RankShift, a systematic study comparing OCR models and vision-language models (VLMs) under controlled variation in dataset composition, matching strategy, and scoring methodology. Using OmniDocBench as a case study, we evaluate the ranking stability of 16 models across 3 dataset configurations, 6 alignment strategies, and 3 scoring metrics. We quantify rank stability using pairwise rank correlation, top-k churn, and pairwise reversal rate. Our results show that the alignment paradigm is the dominant driver of instability: cross- paradigm comparisons between end-to-end and markdown-to-markdown evaluation yield a mean Kendall’s τ ≈ 0.33, with individual pairs as low as τ = 0.23. Dataset composition produces comparable instability across page-type strata (mean τ = 0.48) and layout-type strata (mean τ = 0.49), while the cross-dataset shift from digital to scanned documents is more stable (τ = 0.80). Metric choice is the most stable dimension, with mean τ ≈ 0.78 across NED, BLEU, and METEOR. These findings indicate that leaderboard position is not solely a property of model capability, motivating the need for leaderboards that not only report headline scores, but also the robustness of rankings across reasonable design choices.