One Concept Per Activation: Introspection Methods Fail To Verbalize All Concepts Within Superimposed Activations
Abstract
Introspection methods inject an activation into a language model and decode a natural-language description of it. Published evaluations of these methods score each description against a single target concept, and therefore do not measure whether a description accounts for every concept present in the activation. Because activations are superimposed, a description that names one concept may omit others. We construct activations with known concept shares by mixing pairs of unit-normalized SAE decoder directions, and we measure the share below which a present concept is no longer named. Across four methods, two SAEs, two base models, two training objectives (label supervision and reconstruction), and two tuning scopes (adapters on a frozen model and full fine-tuning), the second concept is named at the rate of an absent concept once its share falls to 25%; capacity and training objective are associated with differences at parity but not with movement of the threshold, and the one method retaining detection at 25% loses it at 10%. This threshold bears on the use of introspection methods for safety auditing, since a concept an auditor is looking for need not be the dominant component of the activation at the position being read, and a description that omits it is indistinguishable from one produced when it is absent. These results indicate that current introspection methods do not yet provide complete accounts of superimposed activations, and that verbalizing non-dominant concepts is a capability they would need before a description can be read as evidence of what an activation does not contain. The constructed-activation protocol applies to any injection-based method, and we make our code, pair sets, and scored descriptions available.