Control Reinforcement Learning: Token-Level Mechanistic Analysis via Learned SAE Feature Steering
Abstract
Sparse autoencoders (SAEs) decompose language model activations into interpretable features, primitives for mechanistic circuit analysis. Yet existing methods identify task-specific circuits via attribution or apply fixed-direction steering, but learn no task-reward-optimized policy for per-token feature amplification. We formulate token-level feature intervention as a POMDP over the LLM's layer-causal computation. Control Reinforcement Learning (CRL) trains a policy in this POMDP that selects an SAE feature to amplify at each generation step, yielding per-token intervention traces that expose which features drive task behavior. Adaptive Feature Masking encourages diverse feature discovery while preserving single-feature attribution. The framework provides branch point tracking (tokens where feature choice changes outcomes), critic trajectory analysis (separating policy from value estimation errors), and layer-wise comparison along the residual stream hierarchy. Cross-task transfer is asymmetric (MMLU features harm GSM8K but not vice versa; HarmBench helps XSTest), consistent with partial task computation overlap rather than dataset-specific shortcuts: task-conditional attribution does not require transfer invariance. On Gemma-2 2B and LLaMA-3.1 8B across MMLU, BBQ, GSM8K, HarmBench, and XSTest, CRL provides per-token intervention traces and serves as an SAE quality diagnostic alongside task-accuracy gains. Learned feature steering thus serves as a layer-specific diagnostic complementing activation-based SAE analysis.