Break It to Align It: Annotation-Free Medical Safety Alignment via Preference Inversion
Abstract
Safety alignment in specialized domains requires preference pairs that contrast safe and plausibly unsafe responses to the same prompt, yet such pairs are scarce for ordinary clinical questions. We propose Adversarial Self-Correction (ASC), which constructs them without domain-specific safety annotations. Starting from one backbone, ASC trains a safety-aligned policy and a preference-inverted sibling on opposite orderings of the same general-domain safety data. The two policies then answer the same unannotated medical prompts under matched generation contexts; responses are ordered by policy provenance and used for a second round of DPO from the safety-aligned model. On BioMistral-7B, ASC reduces unsafe XSTest compliance from 10.5% to 5.0% beyond general safety training and lowers MedSafetyBench harm from 1.31 to 1.23, although safe-prompt compliance also falls. On Llama-3.1-8B-Instruct, which already exhibits low unsafe compliance, Stage 3 is close to neutral on most safety measures. These results suggest that preference inversion can provide a lightweight source of domain-specific contrastive supervision when residual unsafe behavior remains.