Values as Style: Disentangling Values from Semantics with One-Way Mixing for Low-Damage LLM Steering
Jiale Dai ⋅ Hongcan Deng ⋅ Liuxian Ma ⋅ Xiaoke Niu ⋅ Guojie Song
Abstract
Activation steering can change an LLM's behavior without updating its weights, but a direction intended to change safety or value stance often also shifts topic, phrasing, and task content. We study a narrower and directly testable question: can a frozen residual-stream state be equipped with an editable interface that separates semantic content from value framing well enough to reduce this collateral damage? We propose a lightweight dual-code interface with a one-way semantic$\rightarrow$value path. Semantic information may ground value recognition, while stop-gradient gating, topic de-confounding, and swap consistency discourage value supervision from rewriting the semantic factor. At inference time, we edit only the value code through a residual delta update. Across two instruction-tuned backbones, the learned interface improves the steering--damage trade-off over dense and sparse steering baselines: it preserves semantics better, reduces topic leakage and benign false refusals, and remains stable under noise, distribution shift, and recomposer controls.
Chat is not available.
Successful Page Load