A Review of Indirect Prompt Injection and Its Defenses Using Large-Scale Public Red-Teaming Competition Data
Yixiong Hao ⋅ Maxwell Lin ⋅ Mateusz Dziemian ⋅ Xiaohan Fu
Abstract
Indirect prompt injections (IPI) hijack an LLM agent through the content its tools return. Progress on IPI defenses is hard to measure because existing benchmarks are unrepresentative and quickly saturated. We evaluate 12 published defenses against 13k+ unique human-authored IPI attacks and 26k+ successful break trajectories from public red-teaming competitions. We find that reported effectiveness largely fails to generalize to the attack set; four of six classifier defenses fall short of their published precision and recall. Among system-level defenses, Plan-Then-Execute, Code-Then-Execute, and Dual-LLM designs generalize best, with one cutting attack success from 61.9\% to as low as 1.6\%. However, attack reduction co-occurs with lower utility, and all system-level defenses leave residual attack surfaces. We also show that 1) the most effective strategies on today's LLMs forge conversational roles rather than issue instructions, 2) effective strategies have mostly not changed over time, and 3) small open-weight models ($<$40B) are increasingly cost-effective screening proxies against frontier targets. These results suggest the need for evaluating LLM defenses on human generated, up-to-date adversarial data, and for LLM agents deployed in high stakes applications to use an ensemble of defenses.
Chat is not available.
Successful Page Load