AxiomShift: What Accuracy Cannot Reflect in Scientific Reasoning Evaluation
Abstract
Language models are increasingly placed inside scientific workflows, handed premises — a protocol, a measurement, a stipulated model — and expected to reason from them. Accuracy is the instrument we use to certify that they can. We ask what accuracy actually measures, and find it cannot distinguish a model reasoning from the evidence it is given from one reporting what it already believes, because on standard benchmarks the two coincide. We introduce AxiomShift, a 700-question benchmark that aims to separate them: in each of seven closed systems drawn from biology, chemistry, physics, economics and mathematics one axiom is altered, so the real-world answer is verifiably wrong, with a principled abstention set alongside. Across 26 models from 3.9% to 98.4% accuracy, decomposing every error into four mutually exclusive failure modes, we find that as accuracy rises from 10% to 90% the two competence-driven modes are eliminated almost entirely while answers leaking the parametric prior are only halved. Prior leakage, then, is not what is left over when a model reasons badly. It is a different kind of failure, one that higher accuracy does not fix, and it concentrates where the supplied axioms do not determine a unique answer. Accuracy cannot separate it from ordinary miscomputation, so it goes unmeasured, which makes it a validity problem for any evaluation offered as evidence that a model can reason scientifically.