Are You Reading the Right Thing? Rethinking Internal Signals for Prompt-Injection Detection
Abstract
Several recent studies propose detecting prompt injections from a language model’s internal state. But these methods do not measure a single security property. Residual-stream probes, task-drift probes, attention scores, role projections, ex posure probes, and concept readouts may capture injection presence, recogni tion, information routing, or behavioral influence. It is therefore unclear what differences in prompt-injection detection performance reveal about the internal signals these methods read. We study how six representation-reading methods vary under a common structured indirect prompt-injection setup on Qwen3.6 35B-A3B and Gemma-4-26B-A4B-it. We evaluate each method by how well it distinguishes injected prompts from benign controls. We then further analyze each detection method’s ability to detect attacks in cases where they are success ful and in cases where they are not. This helps us better assess what properties each detection method might rely on. Methods trained to detect injection pres ence, task drift, or attack-related concepts generally separate unsuccessful injec tions more strongly from benign inputs, whereas attention diversion, exposure, and Role-confusion based methods separate successful injections more strongly on both models. Attention Tracker provides the strongest consistent successful injection separation, reaching 0.949 AUROC on Qwen and 0.887 on Gemma. These results suggest that signals closer to response formation may better encode information about an attack’s success, and that detection performance can depend strongly on the prompt template. We share the data associated with this study at https://anonymous.4open.science/r/paper-evaluation-data-7877