Sparse Internal Control of Language Models
Abstract
Language models are increasingly used in text generation, decision support, and automated interaction, where their behavior must be controlled in a localized, reversible, and selective way. Existing methods either update shared parameters or constrain generation through external interfaces, leaving open whether frozen models contain compact internal control. We introduce sparse internal control, a framework for steering target behavior by applying inference-time interventions to a small set of internal model nodes. The framework formulates control through local target controllability, where candidate site–direction pairs are characterized by their behavioral effects under target, preservation, and energy constraints. This yields two complementary selectors: Driver-OMP, which selects nodes that sparsely reconstruct a desired behavioral displacement, and Coverage-Driver, which favors nodes whose effects cover prompts stably. We further propose a six-axis control-evaluation protocol measuring reachability, signed reversal, dose response, feedback controllability, off-target preservation, and held-out reuse. Across IOI, MMLU MCQA, and sentiment steering on models from GPT-2 small to Qwen3-4B, sparse driver sets act as executable actuators: they steer target behavior, respond monotonically to dose, support closed-loop control, and preserve unrelated next-token behavior. On refusal steering, the same protocol exposes a boundary case that passes all evaluated axes except strict reachability. Held-out reuse separates prompt-specific controls requiring refitted strengths from population-level controls that transfer as frozen interventions. These results show that model control can move beyond output-side steering: mechanistic internal variables can be selected, certified, and reused as localized control handles for frozen language models.