Introspective Coupling: LMs Learn to Explain Themselves Better Than Their Training Targets
Zifan Carl Guo ⋅ Laura Ruis ⋅ Jacob Andreas ⋅ Belinda Z Li
Abstract
When does training language models (LMs) on explanations yield faithful introspection, rather than superficial imitation? Surprisingly, we find that LMs trained to explain the predictions of similar models frequently produce explanations more faithful to $\textit{their own current behaviors}$ than to those of their training targets. This "introspective'' coupling between the model's explanations and behaviors occurs only when the training target explanations remain sufficiently similar to model behaviors over the course of training. This alignment must be preserved throughout the training process: introspection only emerges when explanation supervision is sufficiently behaviorally compatible with the model as it changes. Finally, we show that introspection generalizes to variants of the training problem: when introspection training is run concurrently with training that shifts a model's behaviors, explanations track those behavioral shifts without requiring updated supervision. This holds across a diverse range of tasks, including sycophancy and refusal, and is robust to label noise. These results suggest that introspection training is a viable component of post-training: explanation labels need not be continually refreshed, and faithfulness extends to regions of input space not explicitly supervised.
Chat is not available.
Successful Page Load