Decomposing One Professional-Framing Pipeline: Which Components Shift LLM Safety Boundaries?
Abstract
Jailbreak attacks that frame harmful queries as professional requests can bypass LLM safety training, but we do not know which ingredients in a professional-framing pipeline actually matter. We introduce a pre-specified 2×2×2 factorial-ablation design for jailbreak decomposition, separating access-driving from depth-driving components, and apply it to one concrete pipeline family, Structured Three-stage Framing (STF): a Persona Adoption turn (S1) that establishes professional identity, a Moral Justification turn (S2) that appeals to harm prevention, and a Linguistic Substitution (S3) of framework terminology for colloquial language. Tested on N = 5,000 trials across three closed frontier models (~31,900 total trials across nine LLMs), the decomposition isolates S1 (Persona Adoption) as the dominant contributor to acceptance (OR = 6.5), S3 (Linguistic Substitution) as the dominant contributor to depth (β = 0.91), and S2 (Moral Justification) showing no detectable positive effect at the stated equivalence bounds (pooled OR = 0.83; TOST at [0.67, 1.50]) — a component ordering we write as S1 > S3 > S2 (Persona > Linguistic Substitution > Moral Justification). This ordering holds across three full-rank factorial replications (defensive system prompt, cybersecurity, social engineering; N ≈ 5,000 each) and nine additional robustness checks. Two defensive targets follow: detect persona claims to block access; restrict output depth when responses use framework terminology. The method is portable beyond STF; the empirical result is specific to one pipeline tested primarily through provider APIs, with open-weight corroborations on local inference (§3).