Activation Steering in a Jet-Tagging Transformer: Physical Concept Directions and Their Limit
Abstract
Mechanistic interpretability methods such as linear probing and activation steering have been studied extensively in language models, but many of the concepts they target do not have precise ground-truth definitions. Models trained instead on scientific data give us a useful way to test these methods because many relevant concepts are quantifiable and directly measurable. Before interpretability can be used to uncover scientific structure that is not already known, it is important to establish that these methods can recover and manipulate quantities we do understand. We study this question in jet tagging, a particle-physics classification task, using a Transformer trained directly on particle-level features. We examine four physical quantities (jet mass, particle multiplicity, jet charge, and track displacement) and find that each is strongly linearly decodable from the model’s internal representations. We further identify corresponding steering directions and show that intervening along them changes both downstream representations and model predictions in physically expected ways. At the same time, strong interventions can push representations beyond physically meaningful values, which is an important limitation for using steering methods as tools for scientific discovery.