Conservative Steering: High-Specificity Local Editing of Associative Memories via Curl-Free Energy Shaping
Lakshya Narula
Abstract
Editing a frozen language model—overriding its next-token prediction after chosen prompts—risks collateral damage: related facts and paraphrases silently change too. We present Conservative Steering (CS): at selected layers, add to the residual stream the negative gradient of a learned energy function built from local Gaussian wells placed at the representations of the edited prompts. The field is exactly zero at initialization, decays as $\rho\,e^{-\rho^2/2}$ in the width-scaled distance $\rho$ to the nearest well (a closed-form interference bound), and composes additively, so no anchor data is needed. On CounterFact edit sets applied simultaneously to frozen GPT-2, CS lands every edit from 16 to 250 edits while neighborhood specificity degrades only gracefully ($0.99 \to 0.91$; zero perplexity drift), whereas LoRA and preservation-designed baselines (anchored-/null-space-LoRA, EWC) damage the neighborhood wherever we test them (specificity $\le 0.26$); the signature replicates on GPT-2-medium and a Mamba backbone. Because wells add without moving existing ones, the same edits can arrive one at a time with no replay: CS retains every earlier edit (retention 1.00, backward transfer 0.00) and revises them in place, whereas sequential LoRA keeps only its last (0.06), and GRACE, the closest lifelong editor, retains every edit but at neighborhood KL 1.3 vs. CS's 0.002. Matched-strength, matched-capacity, and curl-only controls locate the mechanism in local, self-aiming placement rather than curl-freeness per se; against ROME, CS trades paraphrase generalization (0.01 vs. 0.78) for far higher specificity (0.99 vs. 0.65). CS is an enumerative, provable, high-precision local editor—preservation by geometry rather than by rule.
Chat is not available.
Successful Page Load