CyberDualEval: Measuring Dual-Use Cyber Risks in Frontier Language Models
Abstract
Frontier language-model agents are rapidly becoming capable cybersecurity assistants, creating a dual-use cyber risk: the same capabilities that help defenders find, understand, and validate vulnerabilities can also help attackers weaponize them. Existing cyber benchmarks largely measure either raw offensive capability or generic refusal behavior, but they do not evaluate whether models can preserve benign security utility while refusing assistance that materially advances misuse. We introduce CyberDualEval, a benchmark for measuring this boundary on real vulnerabilities. Each task is decomposed into three phases of increasing operationalization: vulnerability analysis, proof-of-concept generation, and exploit execution. To capture the contested middle ground, we annotate each proof-of-concept task with a Task Weaponization Score that measures how much work remains to turn a minimally sufficient demonstration into unauthorized cyber misuse. CyberDualEval evaluates frontier language models with varied agentic scaffolds and reports refusal, defensive analysis accuracy, and optional exploit validation as distinct signals. Our evaluation reveals two complementary failure modes. Many models over-comply, assisting with high-weaponizability PoCs or exploit requests that should be refused. Conversely, recent safety-tuned models often over-refuse, declining low-risk or defensive requests and thereby reducing benign cyber utility. These results show that cyber safety cannot be assessed by capability or refusal rate alone: safe deployment requires calibrated refusal that tracks weaponizability while preserving useful defensive assistance.