Equal Questions, Unequal Allocation: When Cross-Language Cascade Metrics Change Meaning
Abstract
Matched-budget cross-language cascade scores need not compare the same counterfactual selection problem. We evaluate a same-family small/large cascade on semantically parallel questions in five languages. Raw regret is mechanically confounded by model-pair headroom; normalizing it into selection skill does not remove the problem because the random-to-oracle denominator changes branch at a language-specific positive-headroom rate. Forty of 50 preregistered language benchmark–budget cells are post-saturation, and no registered budget places all ten cells in the same pre-saturation regime. We derive an accuracy-only certificate that detects regime mismatch without item-level logs and demonstrate it on published MMLU-ProX results. The registered ratio contrast was imprecise: realized influence SD was 2.203 versus a planned cap of 0.596, implying a minimum detectable effect of 0.185 rather than 0.05. This precision failure does not affect the exact denominator identity or saturation census. Oracle construction—and its confidence-feature overlap with the router—also changes the pooled sign. As a deployment illustration, one English-reference threshold escalates 20.0% of English questions and 72.5% of Swahili questions. The extra escalation buys substantial incremental accuracy, yet final accuracies remain below English; allocation disparities alone therefore do not isolate router-caused service inequality. We recommend headroom-aware rather than rate-equalizing audits.