When Splitting Verdicts from Evidence Fails: A SciFact Audit with Constrained-Decoding Controls
Abstract
A scientific reading tool can return a correct verdict without the evidence that justifies it. We audit joint versus separate verdict and sentence-selection calls, then test whether format enforcement explains the observed gap. On 172 evidence-positive SciFact development claims, splitting reduces complete joint success from 56 to 22 for Qwen2.5-1.5B and from 78 to 10 for Qwen3-1.7B. A fresh, source-disjoint 96-claim panel balances support, contradiction, and natural insufficient evidence, and compares free with constrained decoding for Qwen3 and SmolLM2-1.7B. All constrained outputs satisfy the tested syntax, but complete joint success remains lower for split calls: 28 versus 42 claims for Qwen3, and 2 versus 20 for SmolLM2. The primary decoder-by-interface interaction is +25.0 percentage points [17.7,32.3] for Qwen3 and -6.25 [-14.58,2.08] for SmolLM2, using two-model-adjusted intervals. Thus format enforcement substantially narrows the Qwen3 gap without establishing the same effect across model families. Across both panels we retain 2,872 generations and separately check syntax, semantic accuracy, rationale completeness, exactness, and cost. The contribution is a controlled diagnosis of scientific-reader interfaces, not a new reasoning algorithm or evidence that valid serialization guarantees reliable scientific claims.