AgentAbstain: Do LLM Agents Know When Not to Act?
Abstract
Agent systems based on Large Language Models (LLMs) are increasingly deployed for autonomous tasks, yet existing evaluations mostly focus on task success rather than whether agents know when to abstain. This gap poses real risks: under ambiguity, conflicting constraints, or tool failures, agents may execute unintended and irreversible actions. To close this gap, we present the first systematic evaluation framework for agentic abstention: the ability to recognize when not to act. At its core, AgentAbstain is a paired-task benchmark that defines 8 abstention categories spanning pre-execution reasoning and runtime discovery, with 263 paired tasks across 42 executable sandbox environments, each pairing a should-act task with a should-abstain variant produced by a controlled perturbation to the instruction, tool, or environment state. Scaling such paired evaluations poses two practical challenges: manually authoring diverse tasks is expensive, and static benchmarks risk data contamination as models evolve. To address both, we propose AbstainGen, a fully automated pipeline that synthesizes sandbox environments and generates paired tasks end-to-end, validated by deterministic replay and semantic LLM judges. Its scalable design enables on-demand regeneration of fresh task instances, and human validation rates 96% of generated tasks as well-designed. Evaluating 17 frontier LLMs across 4 agent harnesses, the best model (Gemini 3.1 Pro) achieves only 59.5% paired accuracy (correct on both the act and abstain sides of each paired task). More importantly, abstention capability is largely independent of general task-solving capability, indicating that scaling task-solving alone will not close this gap. We identify failure modes such as Post-hoc Abstention, in which agents execute irreversible actions before recognizing abstention triggers. For instance, an agent may cancel a flight reservation before noticing contradictory rebooking instructions, leaving the user stranded. These findings underscore the need for rigorous abstention evaluation to develop more trustworthy LLM agents.