When Probes Fail to Steer: Decodability and Causal Control in Closed-Loop Reinforcement Learning
Abstract
Probes can reveal what a policy represents, but not whether those representations control behavior. We test this gap in a frozen Rocket League policy by training predictors of future simulator events from internal activations, then editing those activations during live rollouts. The probes decode aerial touches, dribbles, and challenges well offline, but most edits fail in closed loop: they do little, move the wrong label, or perturb the policy too broadly. One narrow exception remains after held-out testing: boosting one SAE feature through its decoder-column direction raises pooled dribble-label frequency from 42.84 to 65.11 labeled steps per 1k environment steps; averaging same-seed deltas with equal seed weight gives lift +19.29, 95% bootstrap interval [+11.43, +27.13], positive lift on 11 of 12 seeds, and same-state KL 0.00466. We propose that prediction is useful for finding candidates, but intervention tests are needed before calling a representation a control handle.