DACE: Diversity-Driven Adversarial Co-Evolution for Robust LLM Safety Alignment
Peng Yu ⋅ Xiaoyu Wen ⋅ Zhida He ⋅ Ziyuan Zhou ⋅ Han Qi ⋅ Shao Zhang ⋅ Qiaosheng Zhang ⋅ Ying Wen ⋅ Chaochao Lu
Abstract
Ensuring robust safety alignment of large language models (LLMs) is increasingly difficult as adversarial attacks evolve and outpace alignment pipelines built on pre-collected data. Recent co-evolutionary frameworks let the defender chase a moving attacker, yet their training dynamics expose two compounding pathologies: \emph{attack strategy collapse}, where the attacker overfits to a narrow set of high-reward rewrites, and \emph{defense adversarial forgetting}, where the defender loses competence against earlier attacks as the attack distribution drifts. Existing remedies are partial: attacker-side diversity rewards score textual novelty rather than strategy novelty, and defender-side replay relies on discrete judge scores that ignore estimation uncertainty, response stochasticity, and defender non-stationarity. We introduce DACE, a diversity-driven adversarial co-evolution framework that targets both pathologies jointly. On the attacker side, DACE couples an explicit $12\times10$ strategy space (risk category $\times$ attack style) with a \emph{normalized marginal coverage gain} reward, providing a bounded, non-vanishing exploration signal at the strategy layer. On the defender side, DACE maintains a unified \emph{Bayesian adversarial replay pool} whose Beta--Bernoulli threat posteriors are refreshed by time decay and sampled via Thompson sampling, anchoring defender training to an evolving threat landscape. Across standard safety, automated-attacker, and general-capability benchmarks, DACE improves robustness to out-of-distribution attacks while preserving general capabilities, and yields broader attacker strategy coverage.
Chat is not available.
Successful Page Load