From Tokens to Tactics: Adversarial Text Optimization in an Axis-Aligned Rhetorical Strategy Space for Harmful Content Detection
Abstract
Harmful-content detectors are increasingly deployed in safety-critical settings, but they remain vulnerable to subtle rhetorical reframings that preserve the underlying claim while altering its presentation. Existing adversarial robustness methods are typically token-level or prompt-driven, hindering mechanism-level attribution and targeted repair. We propose Adaptive Rhetorical Adversarial Optimization (ARAO), which operationalizes rhetorical theory as an axis-aligned strategy space with directional and interpretable controls for adversarial rewriting. A planner-rewriter pipeline converts strategy vectors into non-contradictory rewrites, enabling mechanism-level failure attribution and axis-targeted repair. Experiments on misinformation and extremist-rhetoric benchmarks show that ARAO improves robustness by 2.35 ROC-AUC points on average over the strongest baselines on attacked sets, while achieving the best performance on original content.