Chain-of-Thought Oversight Should Not Treat Faithfulness as Monitorability
Abstract
Chain-of-thought (CoT) oversight is often motivated by a simple premise: if a model exposes its reasoning, then its behavior should be easier to monitor. This position paper argues that the premise conflates several properties that can come apart in practice. To make this separation explicit, we distinguish four properties: reasoning visibility (whether a trace is disclosed), mechanistic faithfulness (whether reasoning step content is supported by internal activations), behavioral monitorability (whether a specified monitor can recover a target property on held-out cases), and mechanistic localizability (whether a compact causal mechanism can be found). Claims about one require evidence specific to that property. We support this position with two mirror-image case studies on open-weight reasoning models. In security vulnerability reasoning, probes over internal activations confirm that key reasoning steps are faithfully represented, yet a reference-assisted LLM jury cannot reliably judge whether the resulting analysis is correct on held-out cases. In authority-hint factual reasoning, the pattern reverses: an external monitor detects the target behavior near-perfectly, but repeated mechanistic analyses fail to isolate a compact causal circuit. These results show that internal support does not guarantee monitorability, and monitorability does not guarantee localizability. We argue that CoT oversight papers should adopt property-specific reporting: every claim should state the reasoning access tier, target property, monitor class, monitor competence, held-out protocol, baseline, and evidence type.