Carving New Reasoning Paths: Activation Steering for Conceptual Exploration
Raphaël Sarfati ⋅ Jiajun Bao ⋅ Gurbir Arora ⋅ Owen Lewis
Abstract
Scientific discovery often begins by considering several plausible explanations. Reasoning LLMs can approach complex problems by generating elaborate chains of thought; yet, upon sampling, reasoning traces tend to differ on the surface while converging to the same conclusion. Here, we use causal interpretability to identify answer-specific directions in internal representations and intervene on them during prompt encoding. These interventions expose alternative conclusions more effectively than ordinary sampling, by nudging a chain of thought to land on a pre-determined conclusion. On multiple-choice question dataset~MMLU with Qwen3-8B, steering increased the rate of assigned alternative answers by $5.42$pp over ordinary sampling. The effect extended across additional reasoning models and transferred to GPQA. These results position activation steering as a method for conceptual exploration upstream of discovery. The method broadens the candidate conclusions available for testing, but novelty and truth still require external examination, such as a supervising critique (human or artificial).
Chat is not available.
Successful Page Load