AxiomShift: Altering Axioms to Evaluate LLM Reasoning at Knowledge Boundaries
Abstract
Mathematical reasoning is a central proving ground for AI. Yet accuracy alone cannot establish whether a model follows supplied premises or parametric knowl- edge, or recognises when those premises are insufficient. These issues concern two knowledge boundaries: internal–external, between supplied and parametric knowledge, and known–unknown, between what the premises determine and what they leave unresolved. Wang et al. [2026] formalise procedural prior leakage under internal–external conflict in SQL and Pandas. We introduce AXIOMSHIFT, a 700-question benchmark examining both boundaries across seven mathematical and scientific worlds. Each question either alters an axiom within a derivation or introduces missing or conflicting premises, with a verifier providing ground-truth answer for the altered system and its counterpart in the prior system. Across 26 models, we analyse four failure types and identify two forms of prior leakage: prior- overriding, where a familiar rule displaces a supplied premise, and prior gap-filling, where prior knowledge supplies a missing premise instead of abstention. Among the five most accurate models, prior-leaking answers account for three quarters of failures to abstain. Across models, prior leakage becomes less frequent as accuracy improves but accounts for a growing fraction of the total errors.