Necessary but Not Sufficient: Auditing Same-Answer Counterfactuals for Chain-of-Thought Faithfulness
Abstract
Chain-of-thought (CoT) monitoring is a leading oversight tool proposed for agentic systems. It assumes a model's written reasoning reveals what actually drove its answer. We audit that assumption in single-turn hinted question answering with an answer-conditioned counterfactual design. Two matched prompts suggest different wrong answers. Among traces where the model adopts the first, we label whether the second would have changed the answer (dependence). Prompt-end linear probes predict this dependence label at AUROC (area under the ROC curve) 0.61-0.65 across three 7-9B models, ahead of or within 0.01 of generic LLM text monitors that never see the note (0.51-0.62). But a text judge that reads the CoT and is shown both candidates, asked the label's own counterfactual question, matches or beats the probe (0.65-0.67), and a battery of audits shows the probe's signal is not evidence of hidden causal provenance. A baseline that reads no hidden state, only candidate plausibility, task metadata, and public text, reaches 0.76-0.84. Adding hidden states to it yields no lift. Probes do just as well on prompts that never contained the hint. Choosing hint candidates the model finds equally plausible collapses the signal. Small-sample interventions produce no reliable change. Under stronger prompt injection, a core agent-security threat, the ordering reverses for the three main models and text monitors win. Same-answer counterfactuals are necessary for evaluating causal disclosure, but not sufficient. We distill the audit into a checklist for probe-versus-monitor claims.