Measuring AI Agents' Progress on Multi-Step Cyber Attack Scenarios
Abstract
We evaluate the autonomous offensive cyber capabilities of frontier AI agents on two purpose-built cyber ranges---a 32-step corporate network attack and a 7-step industrial control system attack targeting a simulated cooling tower---complemented by a private 82-task CTF suite spanning four difficulty tiers to measure narrow cyber skills. Estimated human-expert solve times are roughly 20 hours for the corporate range, 22 hours for the cooling tower, and from 1 minute to 40 hours across the CTF suite. Across twelve models released over a twenty-month period---from GPT-4o in August 2024 to early predeployment checkpoints of Mythos and GPT-5.5---we observe two trends. First, performance increases approximately log-linearly with inference-time compute, with no plateau observed for the strongest models up to 100M tokens. Second, newer frontier models generally complete more steps at fixed token budgets. GPT-5.5 and the Mythos Preview checkpoint each completed all 32 steps in at least one run, the first end-to-end completions of the corporate range; GPT-5.5 averaged 22.3 of 32 steps across ten 100M-token runs. On the control system range, the latest models reached 4 of 7 steps, primarily by directly probing the operational-technology protocol rather than following the intended chain of IT compromise and authenticated control. These results indicate rapid improvements in autonomous cyber capability and suggest that complex but undefended multi-step ranges are no longer reliably distinguishing the strongest frontier models. Future evaluations should emphasise hardened environments, active defences, defensive-control effectiveness, and operational security---priorities that grow more pressing as the threat of AI-driven cyber attacks becomes increasingly realistic.