SAJE: Turning Model Specifications Against the Model
Nirav Diwan ⋅ Xiangwen Wang ⋅ Varun Chandrasekaran ⋅ Gang Wang
Abstract
LLM developers publish specifications that describe intended model behavior. OpenAI's model spec, for instance, lists safety rules, behavioral constraints, and edge-case guidance. In this paper, we investigate whether model specifications give adversaries an advantage in black-box jailbreaking. We operationalize this threat using an adaptive black-box jailbreak search framework, *SAJE: Specification-Assisted Jailbreaking Exploits*, which uses model specifications to initialize and guide the search. Our key insight is that ambiguities within and across model specifications can expose promising regions for jailbreak search. Our evaluation across three target models and four evaluation metrics reveals SAJE substantially outperforms state-of-the-art baselines (by up to +25 pp). Notably, SAJE reaches an ASR $\geq$ 80\% using twenty-five target model queries against OpenAI's open-source reasoning model *gpt-oss-20b* using OpenAI's model spec, while producing detailed harmful information on more than $2\times$ than several state-of-the-art black-box jailbreak baselines. Our findings suggest that model specifications can serve as attacker-side auxiliary information that improves the effectiveness and efficiency of jailbreaks while providing more detailed harmful information. Finally, we briefly evaluate defensive practices, including guardrails, and identify open problems for future work.
Chat is not available.
Successful Page Load