ARCANA: A Benchmark for Abstraction and Analogical Reasoning
Abstract
Analogical reasoning is widely considered a fundamental component of human intelligence. Few benchmarks are designed to explicitly test far analogical reasoning—reasoning between domains that are superficially dissimilar yet share deep relational or structural similarities. In this work, we introduce an analogical reasoning benchmark called ARCANA. The tasks of ARCANA require performing far analogical mappings from real-world objects or physical phenomena to abstract grid-based transformations, different from the well-known ARC-AGI benchmark which primarily emphasized innate human priors. We validate ARCANA through a human study, demonstrating that humans solve these tasks reliably. In contrast, evaluations across a range of large language models show a distinct performance gap compared to humans, revealing systematic failure modes that remain challenging for current models. To address this gap, we propose \textsc{ConceptSIR}, a reasoning framework based on concept slippage, introspection, and reflection. Experiments demonstrate that \textsc{ConceptSIR} empowers standard LLMs to solve ARCANA tasks that were previously unsolvable under baseline evaluation. Overall, our results suggest that far analogical reasoning remains an open problem for contemporary LLMs. All the materials will be available at https://sites.google.com/view/arcanacogtest.