Decomposing how prompting steers behavior
Abstract
Prompting steers large language models (LLMs) and vision--language models (VLMs) without weight updates, but it remains unclear how a change in instruction reshapes internal representations to produce a behavioral effect. We introduce a nested geometric decomposition framework that treats prompting as a transformation of the representational geometry for the content following the prompt. We ask what class of mathematical transformation best explains the effect of prompting by finding the best alignment between representations of the same stimulus set following different prompts. For each prompt pair, we fit a sequence of increasingly expressive stimulus-invariant maps: translation, rigid transformation with uniform-scaling, sequential axis scaling, affine, and nonlinear transformations. We then test these maps causally by replacing a single layer's prompt-A hidden state for a new set of stimuli with its mapped counterpart and measuring recovery of prompt-B representational geometry and behavior. Across three LLMs, three VLMs, and six text or image datasets varying in style, emotion, scene content, and number, prompts consistently reshape representational geometry toward the instructed task structure. In the cross-validated nested variance decomposition, much of the prompt-induced activation change is explained by shape-preserving maps: translation and rigid transformation with uniform-scaling. The tier profiles reveal model- and task-specific routing strategies, differing in how much transformation classes explain variance and where along the layer hierarchy their contributions emerge. Crucially, although translation and rigid transformation tiers already improve behavioral agreement, affine transformation is the first tier to nearly recover target-prompt task geometry and produces corresponding gains in behavioral agreement. This suggests that cross-dimensional linear mixing may be a key contributor to how prompts reorganize representations toward the instructed task structure. Our framework provides a general way to decompose prompt-induced representational change into interpretable geometric components, revealing how a model routes task-relevant structure to produce prompt-driven behavior.