Treatment-Correlated Instrument Error in Post-Fine-Tuning Safety Evaluation: A Single-Family Case Study
Abstract
A common check that fine-tuning has not damaged a model's safety sends the new checkpoint harmful requests under one standard condition and counts refusals with an automatic detector. In our own study this check went wrong in two ways, and both errors depended on which model was tested. First, our refusal detector recognized "I can't" only with a straight apostrophe. The untrained and generically tuned control models occasionally wrote it with a curly one (U+2019); the persona-tuned models never did, so the detector missed refusals only in the controls. Correcting it moved the generic control's result by +0.014 in the main experiment, within noise, and by +0.093 in an earlier grid, where it reversed the sign. Second, a heuristic meant to flag responses cut off by the length limit also flagged responses the model ended mid-sentence on its own. Under a permissive system prompt the persona-tuned models did this far more often (7–28% of responses against at most 1%), so discarding flagged responses removed part of the very difference being measured. With both fixed, the case study (Llama-3.1-8B with LoRA, 313 harmful prompts, three training seeds) finds that persona-tuned models keep refusal high with no system prompt but refuse much less under a permissive one, at the same weights. Measured beyond the untrained model's own drop, this extra fall in refusal is 0.306 and 0.381 for the two persona-tuned models against −0.012 for generic instruction tuning, in the same direction at every seed. A second evaluator, which scores how useful a response is for causing harm, moves the same way, by less. Single-labeller human validation of 483 records from an earlier generation finds two over-firings and one miss. The evidence supports an audit procedure and a bounded single-family case study, not a claim about harm specifically or about other model families.