Benchmarking Long-Horizon Agents with Checkpointed Exploration in NetHack
Jonathan Liu ⋅ Alex Zhang ⋅ Seth Karten
Abstract
Games have proven their value as benchmarks for evaluating agents, driving advances in model training and scaffolding techniques. NetHack remains an unsolved game, despite extensive research employing both classical reinforcement learning and imitation-learning approaches. In this work, we revisit NetHack as an LLM-agent benchmark, asking what progress can be made by modernizing NetHack and its evaluation stack. We refactor the NetHack game engine to allow saving and restoring game state, configurable difficulty, and custom observation and action interfaces. We propose a new metric, $BAL_{\min}$, the smaller of a trace's dungeon-depth and experience percentiles among human players, to measure long-term sustainability. Removing obstacles such as interface friction, inaccessible information, and navigation difficulties improves $BAL_{\min}$ from $1.32$\% to $1.88$\%. Two meta-harness algorithms do better: Continual Harness raises $BAL_{\min}$ to $1.96$\%, and Prime-agent-explore, an orchestrator-based search method, reaches $11.38$\% average $BAL_{\min}$ ($+10.1$\%) and $34.66$\% $BAL_{\max}$ ($+24.7$\%). These gains transfer to NetHack with fog-of-war, where Prime-agent-explore reaches $30.9$\% $BAL_{\max}$ and $6.0$\% $BAL_{\min}$. Although NetHack remains difficult for language-model agents, Prime-agent-explore beats prior LLM results ($6.77$\% $BAL_{\max}$ and $1.84$\% $BAL_{\min}$) by $4.6\times$ and $3.3\times$ respectively under fog-of-war and, for the first time, exceeds the mean and median human $BAL_{\min}$ scores.
Chat is not available.
Successful Page Load