JailBound: A FOL-Guided Jailbreak Evaluation Framework for Revealing Safety Boundaries of LLMs
Fazong Wu ⋅ Ming Yang ⋅ Xin Wang ⋅ Zhenyong Zhang ⋅ Xiaoming Wu
Abstract
Large language models (LLMs) are now deployed in a wide range of real-world applications, making it important to evaluate how reliably they resist jailbreak attacks. However, existing jailbreak evaluations mainly rely on manually collected prompts or discrete text-space optimization, which limits their coverage and makes them difficult to extend to new threat settings. We present $\textbf{JailBound}$, a jailbreak evaluation framework that combines automated benchmark construction with intent-preserving attack optimization in embedding space. JailBound organizes evaluation instances with a hierarchical threat taxonomy spanning risk categories, application domains, and attack types, and uses this structure to generate meta-attack prompts with a fine-tuned meta-attack generator. It further formulates jailbreak evaluation as an embedding-space attack optimization problem and uses a first-order loss (FOL)-guided dual-branch search to jointly identify high-value vulnerable regions and safety boundary states. Under a unified evaluation protocol, we study 46 LLMs from 13 model families. The results show that JailBound covers a broader range of risk settings than existing prompt-based benchmarks, and that its optimized attacks transfer nontrivially across model families while supporting finer-grained analysis of vulnerability patterns and safety boundary behavior. $\textcolor{red}{\text{Warning: this paper includes examples that may be offensive or harmful.}}$
Chat is not available.
Successful Page Load