AnchorRep: Defending LLMs Against Cross-Model Adversarial Transfer via Representation Repulsion
Gal Wertheizer ⋅ Rom Himelstein ⋅ Tomer Peretz ⋅ Avi Mendelson
Abstract
Adversarial attacks optimized on a single open-weight LLM can transfer to and jailbreak architecturally different models, allowing an attacker with white-box access to one model to compromise independently deployed systems. This creates a shared vulnerability across models, yet existing defenses are not designed for this cross-model threat. We find that transfer aligns with shared internal representation geometry, making it a natural defense target. We find that cross-model transfer aligns with shared internal representation geometry, making it a natural defense target. \textbf{AnchorRep} targets this geometry directly with a lightweight LoRA adapter that pushes the defended model's internal representations of harmful prompts away from those of a frozen anchor model on the same prompts. Training uses a small set of harmful prompts and no adversarial examples. Across five models and four architectural families, AnchorRep reduces cross-model attack success rate to $\leq$1.1\% on 2{,}000 transferred attacks (0\% on two), including the largest drop on Mistral ($36\% \to 1.1\%$). Existing defenses can reduce transfer, but only at high cost—either inducing up to 77\% degenerate benign output or increasing over-refusal by up to 18\%. Because such degenerate benign outputs are not captured by standard refusal-based metrics, we introduce the \textbf{Benign Garble Rate} to quantify them. Our results suggest that cross-model robustness can be achieved by shaping representation geometry, without requiring attack-specific training. \smallskip \noindent\faGithub~Code, configs and logs: \href{https://anonymous.4open.science/r/AnchorRep/README.md}{anonymous.4open.science/r/AnchorRep}
Chat is not available.
Successful Page Load