WiREBench: Evaluating AI Agents' Capabilities in Reverse-Engineering Black-Box Applications in the Real World
Abstract
As frontier AI technologies dramatically improve their abilities to perform cybersecurity and reverse-engineering tasks, it is critical to rigorously and realistically track their capabilities in these domains. Increased agent capabilities in reverse-engineering will scale black-box vulnerability discovery, which will have dramatic implications on software security at large. In this work, we present WiREBench, a unique benchmark derived from popular real-world mobile applications, to measure agents' end-to-end vulnerability discovery and reverse-engineering capabilities. This benchmark, containing 22 protocols and 62 subtasks, evaluates agents abilities to identify and exploit historical vulnerabilities in binaries' proprietary network cryptography, via black-box reverse-engineering. To resolve these tasks, agents must chain together static binary analysis, understanding of applied cryptography, and complex tool use via dynamic instrumentation of mobile applications. Beyond benchmarking, we ran WiREBench against real-world applications, identifying and confirming 65 protocol vulnerabilities in applications used by hundreds of millions of users. These results demonstrate WiREBench's utility not just as a benchmark to assess current agent capabilities, but also as a pipeline for auditing cryptographic protocols and reverse-engineering in the real world.