Memory Inception: Latent-Space KV Cache Manipulation for Steering LLMs
Abstract
Steering large language models (LLMs) is usually performed through either instruction prompting or activation steering. Prompting offers strong control, but repeatedly caches guidance tokens and can clutter long interactions. In contrast, activation steering ,while compact, typically offers weaker control and less fine-grained steering because it does not support large structured reminders. We introduce memory inception (MI), a training-free method that steers in latent attention space by inserting text-derived key-value (KV) banks only at selected layers. MI treats steering as selective KV allocation, where reminder content need not occupy the full-layer prompt cache in order to induce implicit behavioral guidance. We formulate MI as attention over prompt, target, reference, and auxiliary banks, with canonical pre-RoPE key storage for reuse across positions and architectures. On matched personality-steering tasks, MI offers a balance between prompting and contrastive activation addition (CAA), approaching or exceeding prompting in raw control strength and CAA in drift mitigation. MI also transfers to structured heuristic guidance on physical reasoning domains, such as the PHYSICS benchmark, where it outperforms visible prompting on average, wins 10/12 subject times mode cells, and cuts content-matched KV storage by up to 118 times. These results position MI as a powerful steering method when guidance is persistent, structured, or expensive to keep in the visible transcript.