Models that Miss Together: Dependence-Aware Conformal Routing for Language Models
Siva Rajesh Kasa ⋅ Karan Gupta
Abstract
Risk-controlled routing selects, per query, a subset of models from a pool so that the probability that every selected model answers incorrectly stays below a target level $\alpha$, while calling as few models as possible; under deployment budgets, fewer calls means directly lower latency, energy, and cost at a stated reliability. The state of the art calibrates a conformal threshold over per-model router scores, although the certified event, that all selected models miss at once, is a joint property of the pool: the joint law of correctness factors into per-model marginals and a dependence structure, and per-model scoring discards the second factor. We make that factor operational. We prove that the all-miss probability of any candidate set is an identified value of the correctness subcopula, that marginal ranking is optimal exactly when failures are conditionally independent, and that any dependence-aware policy frozen on a design split inherits the distribution-free conformal guarantee unchanged. We then replace the per-model score with a sequential coverage score, ordering models per query by how much each reduces the estimated probability that all previously selected models miss; the baseline is recovered as the independence special case. Across candidate pools of 100 to 500 models drawn from a 3{,}811-model leaderboard, the new score reduces mean model calls by 57 to 83 percent at $\alpha=0.02$ at equal certified risk, with the advantage growing in pool size, and a validation-based selector over benchmark-level and query-level policies yields one deployable policy per benchmark. Every estimator is tuned only on its own validation split, and a shuffle control that destroys co-failure structure while preserving marginal accuracy measures how much of each gain survives.
Chat is not available.
Successful Page Load