Reasoning Depth Erodes the Premise Verification That Post Training Builds
Govind Arun ⋅ Shreeya D Lakshminath
Abstract
Whether a language model can tell that a problem does not determine an answer has mostly been measured on finished models, which cannot say when the ability arrives in training or what takes it away. Instella publishes intermediate checkpoints spanning pretraining and post training together with its full data mixture, so both questions become measurable. A premise deletion probe removes exactly one premise the reference solution consumes, giving 3179 variants over 1779 parent problems, and every abstention rate is read against an answerable control on the untouched parents, where declining is an error. Verification is close to absent before post training and is built by it: discrimination runs $+1.84$, $+10.10$, $+38.08$ and $+52.63$ points across stage one, stage two, supervised fine tuning and instruct, while wrong refusals stay below 3.6 percent throughout, and a reworded licence preserves the ladder. An independently built MATH probe reproduces it, and three instruct checkpoints from other families all discriminate. What training builds then declines with reasoning depth. Abstention on underdetermined problems falls $5.22$ points per additional calculator step while the answerable control stays flat at $-0.11$, a difference of $-5.11$ on $[-6.36, -3.82]$, and in percentage points the decay steepens as the checkpoint improves. Under the answer forcing prompt used for benchmark scoring the strongest checkpoint still answers 97.6 percent of underdetermined items. Controlled injection into Instella at a fixed token budget drives verbatim reproduction to 0.946 and returns the memorised calculation chain, at $+5.44$ on $[+1.46, +9.38]$, without returning the answer, at $+1.52$ on $[-2.12, +5.21]$, so memorising the source text does not by itself restore the answer.
Chat is not available.
Successful Page Load