TALES: Text-Adventure Learning Environment Suite
Abstract
While Large Language Models (LLMs) are increasingly deployed as agents to complete complex real world tasks, assessing their baseline agentic capabilities remains challenging. Many existing benchmarks for LLM agents require the model to operate through a structured interface that sit one layer removed from natural language, making it difficult to distinguish failures of interface interpretation from genuine limitations in agentic capability. Because they rely entirely on grounded, natural-language interaction, text-adventure games serve as a prime sandbox for evaluating these capabilities. However, current implementations in popular benchmarks obscure performance through the leakage of privileged information or through scaffolding modifications to the tasks themselves. To solve this, we introduce \ours, a diverse collection of synthetic and human-written text-adventure games in their canonical forms, designed to evaluate baseline agentic competencies in progressively more challenging environments. We present results over a range of LLMs, open- and closed-weights; evaluate when and where current models fail as agents; and perform a qualitative analysis on top agents to identify key failure modes. Despite an impressive showing on synthetic games, even the top LLM-driven agents fail to achieve 25\% on games designed for human enjoyment. Visualization of the experiments can be found at https://github.com/tale-suite/tale-suite-anonymized.