Do Small Language Models Verify What They Can Solve? Synthetic Derivations with Planted Errors
Abstract
A model that can solve a problem is often assumed to find checking a solution easier, since verification only asks for a yes or no rather than the answer itself. We test that assumption on two small instruction-tuned models, Qwen2.5-0.5B and Qwen2.5-1.5B, using arithmetic and modular-arithmetic derivations generated procedurally with one planted error in half of them. On the same 640 problems we measure solving, verifying (is this derivation correct?), and localizing (which step is wrong?). Solving with chain of thought reaches 48% for the small model and 68% for the large one, and collapses without chain of thought (under 6%). Verification does not follow. Asked for a one-word Yes or No, both models say Yes to everything: they accept correct derivations 100% of the time and reject corrupted ones 0% of the time, for a balanced accuracy of 0.50 and an AUC of 0.51 to 0.54. A full chain-of-thought recheck barely helps: the large model reaches 52% balanced verification accuracy against 68% solving, and the small model stays at chance. The result is a solve-verify gap that runs the wrong way. The 1.5B model solves 68% of the problems but both solves and verifies only 6%, and 62% of problems fall in the "solved but not verified" cell. Error localization is weakly above chance (24% from step-number token probabilities against a depth-dependent floor) and is driven by a recency bias: the model guesses that the error is in a late step regardless of where it is, so it does best exactly when the planted error sits in the last quarter of the derivation.