BraveATA: Benchmarking Broad and Verifiable End-to-End Automated Theoretical Analysis of Large Language Models
Abstract
Theoretical analysis plays a crucial role in understanding the mechanisms, capabilities, and limitations of large language models (LLMs). However, existing theoretical studies of LLM remain largely human-driven, making it difficult to systematically explore the rapidly expanding space of LLM-related phenomena and mechanisms. In this paper, we introduce the task of Automated Theoretical Analysis of LLM, which requires LLMs to discover, formalize, and prove theoretical mechanisms underlying LLM behavior. To support this task, we construct BraveATA, the first broad and verifiable benchmark for automated theoretical analysis of LLMs. BraveATA is built from 178 papers covering 11 research directions. For each paper, we manually extract the core claim and annotate the corresponding informal theorem, formal theorem, and proof, with additional Lean4 formalizations to support objective and reproducible evaluation. Experiments on 6 mainstream LLM families show that current models remain far from fully automating theoretical analysis. They struggle to construct faithful mathematical abstractions of LLM-specific concepts and, most critically, to produce compilable formal proofs. We further explore a multi-agent framework for end-to-end theoretical analysis based on BraveATA, which reveals that current LLMs still have substantial difficulty generating core claims that are both novel and valuable. Our benchmark and findings provide a foundation for evaluating and advancing LLMs as automated theoretical researchers for future research.