BoardGameArena: A Multi-Dimensional Benchmark for Strategic Reasoning of LLMs in Board Games
Abstract
Large language models (LLMs) excel at reasoning across diverse domains, from mathematical problem-solving to code generation, yet struggle with strategic reasoning that requires long-term planning, opponent modeling and integrating tactical calculation with strategic understanding. Board games, with their fully observable states, perfect-information environments and objectively verifiable decisions, offer ideal testbeds for probing these capabilities. We introduce Board Game Arena (BGA), a multi-dimensional benchmark that evaluates LLMs' strategic reasoning along three cognitive axes—deductive, inductive, and abductive reasoning—through three challenging task types across five classic board games, yielding 45 sub-datasets with over one million samples. We evaluate 28 state-of-the-art LLMs on BGA and find that current models exhibit fundamental limitations in strategic reasoning for board games. While models demonstrate competence in understanding game rules and identifying legal moves, they struggle significantly with selecting optimal moves and applying strategic principles or tactical analysis. These results position BGA as a challenging and diagnostic benchmark for tracking progress in LLMs' strategic cognition.