CoTs as Probabilistic Programs: A Programmatic View of Thinking Step-by-Step in Language Models
Abstract
Chain-of-thought (CoT) prompting and related strategies that elicit intermediate reasoning traces have fundamentally changed how language models are used, evaluated, and developed. CoT now touches many areas of contemporary research beyond prompting, from advanced test-time inference and model tuning to interpretability, yet much of this work treats CoT traces in task-specific ways, without a shared formal account of what they are or how they compose and support inference. We argue that probabilistic programs — ordinary programs extended with stochastic choices and probabilistic conditioning — provide a natural framework for modeling CoT reasoning. To demonstrate this, we derive a small discrete probabilistic programming language for CoT in which reasoning steps are stochastic choices, branches encode dependencies, trace likelihoods score executions, and answers are return values. Even with a minimal reachability-based semantics, this language enables a first-principles analysis of trace likelihood, exposing why likelihood alone can be a misleading model of reasoning. Motivated by this limitation, we add differentiable operators for composing scores and feedback over CoT traces. The resulting framework links several uses of CoT — scoring, evidence aggregation, feedback propagation, model tuning, and step-level sensitivity analysis — within a single probabilistic program. Through case studies, we show how this approach yields new analysis techniques and test-time inference strategies, offering a promising direction for programming-language-based approaches to language-model reasoning.