A Simple Class-Agnostic Approach to Enhance Fair Adversarial Training
Abstract
AI security has become a critical issue as Deep Neural Networks (DNNs) are inherently vulnerable to adversarial attacks. While previous studies have demonstrated that adversarial training is one of the most effective defensive strategies, improvements in robustness often show significant class-wise disparities. This has established fair adversarial training, which aims to enhance the robustness of the worst-performing classes, as a crucial research direction. Although existing class-wise approaches can improve performance by adjusting individual class weights, they face severs scalability limitation on large-scale datasets due to the sparse volume of data available per class. To mitigate this issue without relying on explicit class-wise information, we propose a novel tri-regime optimization framework that fundamentally decouples the distinct sources of adversarial error at the sample level. Specifically, our method systematically suppresses unlearnable outliers via reliability gating, prioritizes genuine boundary threats through dynamic reweighting, and down-weights already safe examples to prevent redundant optimization. By isolating and prioritizing the effective adversarial frontier, our method strictly prevents robust overfitting and achieves superior worst-class robustness while maintaining high clean accuracy.