Corrupted Plans, Clean Traces: What Planning-Execution Decoupling Reveals About CoT Monitoring
Abstract
Chain-of-thought (CoT) monitoring is increasingly relied upon to detect misbehavior in LLMs, under the implicit assumption that misbehaving actors must reason about their misbehavior, leaving detectable traces in the CoT. We show this assumption fails for a class of attacks we call planning-execution decoupled attacks, in which misbehavior is injected via a corrupted plan before the actor begins reasoning. Using the investigator agent framework of Li et al. [2025], we discover that exposing an actor to a corrupted reasoning plan upstream causes it to reproduce the flawed reasoning in its own CoT, embedding misbehavior into natural-looking reasoning with few suspicious traces. The attack scales to stronger reasoning models and harder tasks, steering them toward misbehavior while evading detection. Evaluating a suite of thinking and non-thinking monitors, we find that thinking monitors substantially outperform non-thinking ones, though even the best miss a meaningful fraction of attacks. Critically, the relationship between monitor thinking budget and detection is not monotonic: extra reasoning sometimes improves detection but can also hurt it when monitors talk themselves into accepting corrupted CoTs as benign. This refines Guan et al. [2025], who show thinking budget generally helps monitoring; we find detection also depends on whether the additional reasoning is directed toward critical evaluation rather than rationalization.