Context-Reconstructed Semantic Smoothing for Multi-Turn Jailbreak Defense
Abstract
As large language models become more ubiquitous, preserving their alignment in multi-turn exchanges rises in importance. Semantic smoothing has proven to be an effective defense against single-turn jailbreaks, but may be insufficient in multi-turn contexts where the harmful request is distributed across various messages. We introduce Context-Reconstructed Semantic Smoothing (CRSS), which captures the user's underlying request from relevant dialogue history and applies semantic smoothing, aggregating model behavior. We compare CRSS to latest-turn semantic smoothing across two distinct model configurations and conversations from SafeDialBench. We find CRSS successfully blocks all harmful attacks it covered, as opposed to latest-turn semantic smoothing at ~47\%, all while maintaining nominal performance. Furthermore, in an ablation test, we record CRSS detecting more harmful records, while increasing mean harmful risk from 0.242 to 0.630. We also observe drawbacks in coverage, attributed to the addition of a semanticity validator. Overall, these results suggest context reconstruction enhances the potency of semantic smoothing for multi-turn jailbreak defense, with trade-offs in coverage.