Endpoint Definitions Can Reverse Conclusions About Additional Reasoning
Abstract
Additional reasoning can change both the expressed answer and whether an evaluator recognizes a valid endpoint. We audit this effect using eight feedback-free continuations branched from each fixed candidate across two reasoning models and four automatically scored tasks. Under the protocol-defined strict endpoint parser, DeepSeek-R1-Distill-Qwen-7B continuation on ARC-Challenge changes accuracy by -16.35 percentage points (family-wise 95% CI [-30.83, -1.67]). A post-hoc, conservative, gold-blind union fallback yields +18.75 points [+1.67, +35.83], reversing both the sign and the multiplicity-corrected conclusion. The fallback examines direct text before any parser-retry suffix and accepts only explicit answer-bearing forms; it covers every strict-valid output and agrees on 99.8–100% of them. It recovers 371 of 645 strict-invalid ARC endpoints, of which 337 match the task answer. The effect is cohort-specific: DeepSeek GSM8K moves from -11.15 to -5.83 points, whose family-wise interval [-13.33, +0.83] includes zero. These results show that endpoint definition can materially change the measured value of continuation even when alternative parsers nearly agree on their shared domain. We additionally compare continuation with independent restart under similar generated-output budgets in two focal cohorts.