Separating instructions from data as an attention problem
Abstract
We introduce TrustAttention, a defense against prompt injection that learns which attention heads are involved in instruction following and redirects their attention away from untrusted tokens. Other heads remain free to process those tokens, allowing the model to read and reason over external content without treating it as instructions. TrustAttention thereby enforces instruction--data separation without broadly changing the model's behavior. We evaluate TrustAttention on Llama-3.3-70B-Instruct and GPT-OSS-20B against SecAlign++, a state-of-the-art alignment-based defense. Compared with SecAlign, TrustAttention offers greater robustness to adaptive attacks and better utility on out-of-distribution data, while matching or improving security--utility trade-offs on standard benchmarks. Under adaptive attacks, RL-Hammer achieves 97% attack success against undefended GPT-OSS-20B and 66% against SecAlign, but only 9% against TrustAttention. White-box GCG yields 11% against undefended Llama-3.3-70B, 3% against Meta-SecAlign-70B, and 0% against TrustAttention. TrustAttention also avoids the utility failures we observe for SecAlign on few-shot and multi-turn prompts outside the training distribution. These results show that instruction--data separation can be enforced through targeted interventions on attention rather than broad fine-tuning of model behavior.