Red Teaming is a Measurement Problem: Item Response Theory for Calibrated, Cost-Efficient Robustness Evaluation
Sunil Kothari ⋅ Rajiv Lal
Abstract
Red teaming produces a matrix of (attack, model, behavior) success/failure verdicts, but the field lacks a principled way to turn that matrix into a $\it{calibrated}$, $\it{comparable}$, and $\it{cheap-to-collect}$ robustness measurement. We show that Item Response Theory (IRT)---the psychometric framework behind adaptive standardized tests---is the natural tool. Re-analyzing the public $\it{HarmBench}$ results (15 text attacks $\times$ 28 target models $\times$ 403 behaviors; 124${,}$677 verdicts) at zero additional compute, we (i) calibrate a three-way IRT model that places attacks and models on common scales and, held out, predicts verdicts at AUC~0.89 vs.\ 0.75 for a strong baseline; (ii) show that attack $\it{strength}$ and attack $\it{discrimination}$ diverge---the strongest attack (GCG) is among the $\it{least}$ informative for ranking models, while cheap direct requests are among the most; (iii) find robustness to be $\sim$\PCone\%\ unidimensional except for a distinct gradient-attack vulnerability axis. We then split the adaptive-testing promise into two hypotheses---H1 (a short probe set suffices) and H2 (adaptive $\it{selection}$ is what makes it work)---and find H1 holds while naive H2 is a diagnosed $\it{negative}$: Fisher-information selection does not beat random probing, because the attack surface is multidimensional. Two reframed objectives conditionally recover H2: a $\it{cost-aware}$ objective certifies the robustness ranking $\sim$two orders of magnitude cheaper in attacker effort than random, and a $\it{multidimensional}$ CAT recovers the gradient-vulnerability axis that random probing misses---validated on held-out behaviors. All claims carry bootstrap confidence intervals or held-out tests, and the pipeline is fully reproducible.
Chat is not available.
Successful Page Load