Drop the Act: Probe-Filtered RL for Faithful Chain-of-Thought Reasoning
Abstract
Reasoning models can keep reasoning even when they can already give the correct answer. We call these additional, post-commitment steps reasoning theater. We introduce ProFIL (Probe-Filtered Reinforcement Learning), a drop-in extension to Group Relative Policy Optimization (GRPO). Prefix counterfactuals identify post-commitment steps; a lightweight probe is trained once on activations of a base model and then frozen. During RL, the current policy produces rollouts, while a separate frozen copy of the base model teacher-forces the same rollout text to supply the probe activations. High-theater rollouts receive zero reward and zero policy advantage. Across GSM8K, LiveCodeBench, ToolUse, and MMLU-Redux and two model architectures, ProFIL reduces post-commitment theater by 11–100%, raises faithful fraction (including +24pp on LiveCodeBench under an independent GPT-4.1 judge), and shortens chains by 4–19% in three of four domains, while preserving or improving the reported task-accuracy metric. A matched length-penalty baseline worsens theater, isolating commitment detection from generic compression. Frozen-probe audits, an independent pairwise judge, and inference-time steering controls further support the result and address the RL-obfuscation concern. Probe weights, training configurations, evaluation code, and rollout caches are released for all four domains.