TraceGuard: Defending Multi-Turn Jailbreak Attacks via Prompt-Response Risk Signal Tracking
Abstract
Large Language Models (LLMs) have made remarkable progress and are increasingly deployed as black-box API services. However, they remain vulnerable to jailbreak attacks that bypass safeguards and elicit harmful content, especially in multi-turn settings where harmful intent is gradually constructed over the interaction. Existing defenses primarily target single-turn scenarios or require access to model internals, making them insufficient against black-box multi-turn jailbreak attacks with weak single-turn discriminability, cross-turn dependency, and dynamic attack patterns. To bridge this gap, we first analyze harmful intent in multi-turn jailbreak interactions through representation-space risk signals inspired by the Linear Representation Hypothesis. Our analysis reveals that these signals are fragmented across prompts, responses, and turns, progressively evolve over the interaction, and exhibit attack-dependent temporal patterns. Motivated by this, we propose TraceGuard, an online black-box defense framework that detects multi-turn jailbreak attacks by tracking prompt and response risk signals as the interaction unfolds. Specifically, TraceGuard operates in two modes: TraceGuard-S captures cross-view risk representations and their cross-turn evolution, while TraceGuard-A further enables distribution-aware few-shot adaptation to dynamic attack patterns. Extensive experiments across multiple benchmarks and target LLMs demonstrate that TraceGuard achieves superior multi-turn jailbreak detection while preserving utility, maintaining efficiency, and generalizing to unseen attacks.