Model Cascades with Provable Per-Class Quality
Abstract
Model cascades reduce inference cost by routing inputs through cheap models and escalating only when needed, from lightweight image classifiers to small language models. Existing exit rules (confidence thresholds, learned scorers, agreement signals) all optimize for aggregate quality, but none controls quality at the class level: a model that is systematically unreliable on a rare, fine-grained, or minority class silently violates the quality contract, regardless of its overall confidence. We introduce NOMAD, built on a simple insight: the predicted class label, not the confidence score alone, is the correct signal for cascade exit decisions. For each model, NOMAD identifies from held-out data the classes it handles reliably enough to serve as the final answer; a chain-safety certificate verifies that any sequence of models preserves per-class quality end-to-end, not just at each stage in isolation; and a greedy selector routes each input through the cheapest viable sequence, with a provable 4-approximation on cost. On 16 datasets spanning tabular, fine-grained vision, and text (including LLM cascades with 15–20× cost ratios between models), NOMAD achieves up to 40× speedup (geometric mean 4× over the role model) and is the only method among 11 baselines with zero per-class violations on every dataset. Code, configurations, fold indices, and an interactive results explorer (nomad-cascades.surge.sh) accompany this submission as supplementary material for reproducibility.