Depth, Not Breadth: Best-of-N Jailbreaking Beyond Surface Noise
Abstract
Best-of-N jailbreaking spends a query budget on surface variation, scrambling and recasing a request until one draw lands. Input normalization is built to strip exactly that. We ask what a budget buys when its variance is moved into a structural channel instead, holding the search identical across both arms so the encoding is the only difference. Against SAGE, the strongest published self-check defense, best-of-N over a code-completion encoding reaches 67, 22 and 15% of behaviors on three open-weight targets, where that encoding fired once reaches at most 4.7% and the published character search at full budget at most 3.0%: 9 to 75 times the sum of the parts, with bootstrap intervals clearing both ingredients on every target. A 2×2 holding encoding and variation apart shows the two defense families fail to different factors: a transform defense is broken by the depth of the encoding (7→67 behaviors at fixed variation) and a gate by the breadth of the variation (13→57 at fixed encoding). Repeated sampling also inflates apparent robustness, because an attacker who may try N times experiences the maximum over draws while safety results are reported as means: on one target SAGE blocks 99.8% of individual draws yet loses 12 behaviors to a repeat attacker where a classifier gate blocking 95.6% loses 10. The prescription is architectural, not free: do not fuse screening with generation. We measure its attack-side effect only.