Taming CoT Obfuscation in VLMs: From Mechanistic Evidence to Activation-Level Enforcement
Abstract
Reinforcement learning (RL) has substantially improved the reasoning capabilities of vision-language models (VLMs), yet it often triggers chain-of-thought (CoT) obfuscation, a regime where models achieve high accuracy while producing reasoning traces that are ungrounded and difficult to monitor. While prior work documents this decay at the behavioral level, the underlying mechanistic drivers of such representational drift remain poorly understood. In this work, we identify that obfuscation is a representational pathology, where RL-induced optimization causes task-agnostic template features to replace visually grounded content in the model’s activation space. Based on this insight, we propose TAME, an activation-level intervention framework that leverages Sparse Autoencoders (SAEs) to regularize VLM internal features during RL training. Specifically, TAME utilizes an LLM-as-judge monitor to detect behavioral obfuscation and maps these signals to specific internal directions, penalizing template intrusion while preserving visual evidence. By directly intervening on the features responsible for obfuscation, our method restores the transparency and visual grounding of reasoning traces. Experiments on the VIRL-39k and SPA-VL benchmarks across two model families show that TAME improves CoT monitorability (by up to +30.9% on VIRL and +16.7% on SPA-VL over GRPO), while preserving task accuracy and general-capability performance across four standard VLM benchmarks.