Reliability-Aware Metacognitive Routing for Cost-Efficient Mathematical Reasoning
Abstract
Efficient mathematical reasoning with large language models requires balancing answer quality against inference cost. Existing routing approaches often rely on a model's own confidence to decide whether a query should be escalated to a larger model, but generation confidence is not necessarily a reliable indicator of correctness. We propose Metacognitive Federated Reasoning (MFR), a reliability-aware, sequential routing framework for heterogeneous language models in a Client--Edge--Cloud hierarchy. MFR estimates each model's empirical reliability from a held-out calibration set and combines it with a value-of-information rule that determines whether to stop, consult another model, or escalate to the cloud. On GSM8K with open-weight models ranging from 1.1B to 22B parameters, MFR achieves 68.0\% accuracy with 13\% cloud utilization and 11.49\,s average latency, compared with 38.0\% accuracy for a similarly priced confidence-based router. While MFR remains below always-cloud accuracy, it provides a substantially better accuracy--cost tradeoff and highlights the limitations of raw confidence as an escalation signal for heterogeneous mathematical reasoning.