Understanding the Out-of-Distribution Generalization of Chain-of-Thought Reasoning in LLMs
Nuojing Liang ⋅ Xiaotong Yuan ⋅ Pan Zhou
Abstract
Chain-of-Thought (CoT) reasoning substantially improves Large Language Models (LLMs), yet its out-of-distribution (OOD) generalization mechanism remains theoretically underexplored. We study CoT as compositional OOD generalization: the target task lacks complete reasoning trajectories or direct input-output pairs, while training provides only in-distribution (ID) atomic step data. We show that CoT generalizes by repeatedly applying a shared single-step predictor to compose reusable ID transitions. Theoretically, the OOD task risk is controlled by the sum of ID subtask risks, with an error-propagation coefficient independent of reasoning depth/steps $T$. A Rademacher-complexity analysis further gives a generalization gap of $\mathcal{O}(\sqrt{T/n})$, rather than linear in $T$, due to predictor sharing, where $n$ denotes training samples per subtask. For autoregressive Next-Token Prediction, we derive an explicit bound $\mathcal{O}\big(k\big(\sqrt{{T}/{mn}}+\sqrt{{T}/{n}}+{T}/{m}\big)\big)$, where $m$ and $k$ denote context length and output token number per step. Experiments on Atom-Task and GSM8K support the theory, showing that CoT enables short-to-long and compositional generalization, while explicit state tracking is crucial.
Chat is not available.
Successful Page Load