The Off-Switch Failure: When Safety-Repair Evaluation Rewards Model Collapse
Abstract
Fine-tuning on benign data can erode language-model safety, motivating cheap, training-free repairs. We show that an aggressive repair can appear successful by collapsing the model into repetitive, content-free refusals: attack-success rate (ASR) falls toward zero while valid generation is destroyed. We call this failure an off-switch. Refusal-based evaluators treat coherent refusals and degenerate non-answers alike, so minimizing ASR can select a broken model over a genuinely repaired one. We introduce a repair-evaluation protocol built on a three-way response taxonomy and a Composite Failure Rate (CFR) that scores invalid output as a failure without calling it a successful attack. A repair must clear a gate battery on harmful compliance, validity, and capability before deployment is recommended. The failure reproduces across model families and repair strengths, survives alternative decoding and quantization, and evades a deployment-time runtime-monitor evaluation. A naive selection procedure on a held-out split can return it as an ASR-optimal candidate rather than merely miss it, while constrained repair methods stay coherent throughout. Existing jailbreak scoring correctly rates the resulting non-answer as unsuccessful, but that question is narrower than whether the model still works. We map the protocol onto EU AI Act requirements and RBI governance guidance, without claiming a legal certificate, and release it as a readiness assessment designed to produce reproducible audit dossiers.