CryptanalysisBench: Can LLMs do cryptanalysis?
Lukas Fluri ⋅ Avital Shafran ⋅ Nicholas Carlini ⋅ Matthew Jagielski ⋅ Milad Nasr ⋅ Orr Dunkelman ⋅ Eyal Ronen ⋅ Florian Tramer
Abstract
Cryptanalysis---the task of finding attacks against cryptographic schemes---sits at the intersection of mathematical reasoning and programming, two areas where LLMs have made rapid progress. Just like math and programming, cryptanalytic attacks are formally defined and can be unambiguously verified. This raises the question if cryptanalysis might experience a similar acceleration in progress through LLMs. In this paper we introduce $\texttt{CryptanalysisBench}$, a benchmark of 113 tasks spanning six families of cryptographic primitives (block ciphers, hash functions, etc) drawn primarily from four NIST standardization competitions. Each task asks an agent to break an implementation of a cryptographic primitive by winning a formal security game. The benchmark has three tiers: (1) primitives with known practical breaks; (2) scaled-down variants of primitives without one; (3) a challenge set of unbroken production primitives. We evaluate Claude Opus 4.7 and GPT-5.5: both solve a majority of Tier 1 (71.4\% and 78.5\%, respectively), while Tier 2 remains largely out of reach. For Tier 1 successes, we provide a fine-grained analysis distinguishing paper recall from source-level rediscovery. We release $\texttt{CryptanalysisBench}$ both as a forecasting tool to track when AI cryptanalytic capability becomes a serious factor, and as scaffolding for subjecting candidate schemes to more attacks before they are deployed.
Chat is not available.
Successful Page Load