RobustStress: Stress-Testing AI-Generated Text Detectors under Gradual Perturbations
Abstract
Detecting text generated by large language models (LLMs) is a critical line of defense for ensuring information content security. Although recent studies have gradually shifted to evaluating the robustness of detectors, existing benchmarks generally suffer from two primary limitations: (1) they lack granularity in controlling perturbation intensity, making it difficult to systematically evaluate the dynamic changes in detector performance as attack intensity increases; and (2) they fail to investigate the impact of perturbation intensity on detector robustness within complex scenarios, such as cross-domain and cross-model settings. To address these challenges, we propose \textbf{RobustStress}, a novel stress-testing framework for intensity-controlled perturbations. We achieve quantitative classification of perturbation intensity by varying the proportion of synonym replacement and employing graded LLM-based paraphrasing attacks. Experiments reveal distinct patterns of performance degradation as perturbation intensity increases, and further capture the synergistic weakening effect between cross-scenario shifts and adversarial perturbations. RobustStress provides an effective stress-testing framework for evaluating detectors in real-world scenarios, encouraging the community to move beyond performance on clean benchmarks to the development of more robust defense mechanisms against adversarial attacks.