Your Judge Is a Confound: Evaluator and Attack Choice Distort Jailbreak-Safety Measurement for Fine-Tuned LLMs
Abstract
Red-teaming reports for fine-tuned LLMs headline a single attack-success rate (ASR), implicitly treating it as a property of the model. We show it is governed as much by how it is measured. Across 25 configurations (five open-weight families x Base FP16, 4-bit, LoRA, QLoRA, and full fine-tuning), attacked with black-box, white-box, and one replayed multi-turn jailbreak over HarmBench and JailbreakBench and scored by automated evaluators calibrated against 750 blind annotations (412-sample clean consensus), the judge itself is a first-class confound. Evaluator choice inflates measured black-box Delta-ASR roughly threefold and reorders which methods look most harmful: against human consensus, GPT-4o-mini runs at a 45.7% false-positive rate (still 39.2% under a StrongREJECT rubric) versus 8.6% for a fine-tuned classifier. The confound survives judge scale (a frontier judge, GPT-5.6, still runs a 21.5-29.9% FPR), extends to the over-refusal utility axis (a judge swap shifts levels +8.3 pp, ordering preserved), and recurs within a single judge (test-retest self-disagreement on 3-15% of identical completions). Attack choice compounds the problem: the black-box suite and a gradient-guided white-box attack give directionally opposite verdicts for the same method (p = 1.6e-7), and an apparent LoRA "protection" is seed noise (-30.0 +/- 24.1 pp). Only once these are controlled does a model-level signal emerge: a seed-controlled TOST shows LoRA and full fine-tuning degrade black-box safety equivalently (+/- 5 pp) in this benign one-epoch regime. Safety claims about fine-tuned LLMs must jointly report adaptation method, attack family, seed/recipe, and a human-calibrated evaluator, and we release a protocol for doing so.