Boule or Baguette? A Study on Task Topology, Length Generalization, and the Benefit of Reasoning Trace
Abstract
Recent years have witnessed meteoric progress in reasoning models: neural networks that generate intermediate reasoning traces (RTs) before producing a final output. Despite the rapid advancement, our understanding of how RTs support reasoning, and where this paradigm fails, remain incomplete. To promote greater clarity, we introduce PITA: a novel large-scale dataset of over 23 million statements in propositional logic and their corresponding proofs. As a benchmark for robust reasoning, we focus on length generalization: if a model is trained to determine truth or falsity on statements with proofs up to fixed length, how well does it generalize to statements requiring longer proofs? We propose notions of (1) task depth and (2) task breadth, which measure respectively (1) the number of proof steps required to solve a proposition and (2) the number of unique propositions within a task family. We vary these quantities across subsets of PITA, and find that RT models generalize well on relatively broad and shallow subsets, while deteriorating on relatively narrow and deep subsets compared to non-RT baselines. As a controlled point of comparison, we study a separate, tractable transitive inference task that exhibits qualitatively similar behavior. Our accompanying theory explains the scalings observed in this simpler setting, suggesting one mechanism by which breadth can favor RT models while depth can expose long-context weaknesses. Our findings suggest salient benefits and limitations of proof-like reasoning traces in controlled length-generalization settings.