Spelling Drift Breaks Keyword Refusal Detectors
Sarthak Rawat ⋅ Vivek Chauhan
Abstract
Safety evaluations often count refusals with keyword lists that search a model's output for phrases such as “I can't”. We show that these counts can fail in a way that the usual validation does not catch. On a pre-registered grid of 77 conditions that combine weight quantization with activation steering on Llama-3.1-8B, both manipulations change how the model spells its refusals: it increasingly writes “I can’t” with a curly apostrophe instead of the straight ASCII one, and together the two produce far more of this drift than the sum of their separate effects. Keyword lists written with ASCII apostrophes read these refusals as compliance, so their error depends on the very conditions being compared. Under mild steering, where a 4-bit model (GPTQ-4) still behaves normally, the HarmBench list misses 41 refusals on 300 harmful prompts and reports the steering effect with the wrong sign; each missed response is one the same list accepts once the straight apostrophe is restored. Under stronger steering, where the model also refuses most harmless prompts, it misses 163 of 300 refusals, all but one for the same reason. Every sign reversal large enough to matter occurs in this one checkpoint. Yet the usual check does not see the problem: pooled over all conditions, our own pre-registered detector agrees with a corrected version of itself at Cohen's $\kappa = 0.81$, above the conventional 0.8 bar, but at only 0.19 in the worst condition that passes our screens. A single pooled agreement score cannot detect judge error that varies with the experimental condition, the failure that Chouldechova et al. [2026] warn about. We recommend reporting detector agreement per condition and reading it against the size of the effect claimed.
Chat is not available.
Successful Page Load