Chain-of-Thought Is Not Explainability
Abstract
Chains‑of‑thought (CoT) allow language models to verbalise multi‑step rationales before producing their final answer. While this technique often boosts task performance and offers an impression of transparency into the model's reasoning, this position paper argues that rationales generated by current CoT techniques can be misleading and are neither necessary nor sufficient for trustworthy interpretability. We analyze faithfulness as whether CoTs are not only human-interpretable, but also reflect the model’s underlying reasoning well enough to support responsible use. By synthesizing evidence from prior work, we show that verbalised chains are frequently unfaithful, diverging from the true hidden computations that drive a model's predictions. As CoT is increasingly relied upon in application domains, including high-stakes ones such as medicine, law, and autonomous systems, we further argue that this unfaithfulness is not taken seriously enough in these settings: our analysis of 1,000 recent CoT-centric papers finds that approximately 25\% explicitly treat CoT as an interpretability technique, including papers in high-stakes domains which heavily hinge on such interpretability claims. Building on prior work, we make three proposals: (i) avoid treating CoT as being sufficient for interpretability without additional verification, while continuing to use CoT for its communicative benefits, (ii) adopt rigorous methods that assess faithfulness for downstream decision-making, and (iii) develop causal validation methods to ground explanations in model internals.