Answering At Any Cost: Frontier LLMs Are Consequence-Insensitive
Abstract
LLM-based systems are increasingly deployed in domains where incorrect outputs carry real costs, often necessitating expensive human review and verification. We argue that a central issue is not simply error, but consequence-insensitivity: models fail to adjust their behavior according to the cost of being wrong. We study this failure across both explicit utility framings as well as natural-language descriptions of stakes that mirror real-world deployment scenarios. Our evaluation spans agentic coding and mathematical reasoning, covering five frontier model families. Across settings, models systematically under-abstain: they continue to answer or submit patches even when incorrect answers carry significant consequences. The pathology is striking: models continue to submit answers in trivial settings where abstention is strictly dominant, and even when told an incorrect answer will cause nuclear extinction. Further, we find that this behavior is orthogonal to existing benchmarks; increasing model size or capability within a family does not improve sensitivity. Comparing base and instruction-tuned models localizes much of this anti-abstention bias to post-training: base models abstain far more often, though they are not themselves reliably consequence-aware. Finally, we evaluate prior post-hoc interventions alongside in-context learning and fine-tuning, finding that none robustly resolves the failure. Our results suggest that current post-training and evaluation pipelines optimize models to answer, not to act under stakes -- a significant bottleneck for trustworthy autonomy.