Do You CARE to Generalize? Extracting Robust Concept Directions from LLMs
Abstract
Transformer-based large language models (LLMs) encode many high-level concepts as linear directions in the latent activation space. Once identified, these directions support both measurement, the quantification of a concept's presence, and intervention, the steering of the model's behavior. In practice, however, a direction learned from one context often fails when applied to prompts from a new context. This poor transfer can have two distinct sources: spurious correlation between the concept label and dataset-specific features, and genuine heterogeneity in how the concept is encoded across contexts. Methods from out-of-distribution (OOD) literature can address the first, but worsen the second by discarding context-specific structure that may carry concept signal. We introduce Context-Aided Representation Extraction (CARE), which bridges OOD generalization and mechanistic interpretability. Given activations labeled with a concept and an environment, CARE jointly learns a shared direction, optimized to be invariant across environments, and orthogonal environment-specific residuals that capture how the concept varies by context. We evaluate CARE on subject--verb agreement, refusal of harmful prompts, and toxicity. CARE produces directions that measure concepts more reliably under distribution shift and in unseen environments and datasets than existing methods while remaining effective for intervention. Its decomposition also supports a swap-cost diagnostic that identifies when context-specific structure carries concept-relevant signal.