JEDI: Real-Time Jailbreak Defense for LLMs via In-Generation Detection and Intervention
Abstract
Large Language Models remain vulnerable to jailbreak attacks despite extensive efforts to ensure safety alignment. Streaming generation scenarios exacerbate this vulnerability by exposing harmful tokens to users immediately upon generation, a phenomenon known as prefix exposure. Existing defenses often fail to address this real-time constraint or impose substantial latency that degrades user experience. To bridge this gap, we propose JEDI, a real-time defense method that secures streaming outputs via in-generation detection and intervention. JEDI leverages representation engineering to monitor the LLM’s internal semantic drift, using a Cumulative Sum algorithm to identify persistent, harmful intent before it manifests in the output. Upon detecting a risk, the system dynamically injects a steering vector to redirect the generation trajectory toward a safe subspace. Extensive evaluations across 6 distinct models and 11 attack vectors demonstrate that JEDI achieves defense success rates ranging from 93.6\% to 99.4\%, outperforming state-of-the-art baselines. Furthermore, JEDI preserves model utility with a Time To First Token overhead of approximately 0.003 seconds, validating its viability for latency-sensitive applications. Our code and experimental results are available at: https://anonymous.4open.science/r/JEDI-93F9