Store or Pointer: Two Mechanisms for Chain-of-Thought Use
Abstract
Chain-of-Thought (CoT) improves the performance of Large Language Models and has been claimed as a window into their reasoning, yet recent work questions the idea that CoT reflects the computation that produces an answer. Therefore, the nature of the CoT relation with the answer, and its underlying mechanisms, remain obscure. In this work, we ask two questions: first, whether the CoT is necessary to generate the answer and sufficient by itself; second, how models use the CoT mechanistically. To study these questions, we designed a controlled testbed: a symbolic multi-hop task where the CoT follows a path to a terminal reference whose value is stored in a lookup table. We train 81 models across 27 architecture configurations, varying depth, width, and random initialization. Using activation patching, we test CoT necessity by replacing the trace while fixing the prompt, and sufficiency by replacing the prompt while fixing the trace. CoT is necessary in every model: replacing it collapses the probability of the expected answer. Its sufficiency, however, varies substantially. Using complementary interventions, linear probes, and the answers produced under intervention, we were able to distinguish two mechanisms of CoT use: CoT hidden states either (i) store the resolved answer, or (ii) act as a pointer to the lookup table, with retrieval occurring only at answer time. Thus, identical CoT traces can support fundamentally different computations. Which mechanism is implemented is constrained by model size: the pointer mechanism emerges only in sufficiently deep and wide models. Still, size alone does not determine the type of mechanism that will be learned: models with identical architectures and training data can converge to different mechanisms purely as a consequence of random initialization, so a mechanistic account derived from one trained model may not generalize to another with the same architecture and data.