Validity-Gated Generation of Overthinking-Inducing Questions
Abstract
Overthinking benchmarks perturb questions until reasoning models produce long traces, and score severity automatically. We show that much of what such pipelines measure is the question’s fault: across 224 perturbation outputs labeled blind by three annotators, 44.3% of items rated severely overthought are ill-posed, wrongly answered or textually broken (95% CI 34.7–54.1%), and mean overthinking estimates fall by 15–22% when they are removed while medians barely move. We build an admission gate (prefilters, a cross-family coherence veto, a two-judge sufficient-prefix severity measure) and validate it at pre-registered thresholds: precision against blind human validity is 0.746 (CI 0.627–0.837) and its rejects are non-clean three times in four. We then use the gate as a reward mask for an RL question generator and ablate it. Without the gate, GRPO converges within 50 steps to copying the seed verbatim, because an unmodified question is always answerable and still earns prefix-excess reward; with it, the policy learns to rewrite, admitted rewrites are human-valid at the rate of a verified clean set, and on held-out seeds the gated policy leads the ungated one by +7.4 points under a common gate (CI 2.1–12.6). We report the criteria that failed beside those that passed: five-way validity agreement of 0.577 against a 0.6 target, and a gate that leaks a growing class of structural defects the reward pays for.