When the Calculator Lies: Mathematical Agents Defer to Faulty Tools Even When They Know Better
Abstract
Mathematical agents delegate arithmetic to tools, so their reliability depends on what happens when a tool returns a wrong value without any error. We introduce FaultMath, a fault-injection testbed that corrupts exactly one tool result inside a multi-step solution. Faults range from a sign flip to a ±1% change, and because every intermediate quantity is known, each final answer can be classified by its numeric provenance as resisted, deferred or derailed. We report a Deference Index (DI), the probability that a model adopts the corrupted value on problems it solves correctly without any tool; this conditioning is what "knows better" means in the title. Across 99k episodes and nine open models (Qwen2.5-Instruct 0.5–14B, Qwen2.5-Coder, R1-Distill, Llama-3.2), competent models adopt plausible wrong values 77–100% of the time under a neutral prompt, usually without comment, and still do so 26–54% of the time when told they may ignore the tool. Deference rises steeply with fault plausibility. Larger models flag implausible faults much more often, but plausible-fault deference never falls below 86% between 1.5B and 14B. A real function-calling loop is at least as deferential as a static log at 3B and 7B and does not respond to verification prompts there. An explicit verification instruction cuts deference by up to a factor of three, at three to five times the token cost and, for the general instruct models at 1.5–7B, with a loss of accuracy on correct logs.