LLM-as-Environment: Benchmarking Open-Ended Decision Making with Language-Model Simulators
Caroline Choi ⋅ Neil Band ⋅ Nick Rui ⋅ Sandra Yang ⋅ Zeyneb Kaya ⋅ Ludwig Schmidt ⋅ Tengyu Ma
Abstract
Language models increasingly make consequential decisions over long horizons, yet most benchmarks assume environments whose states, actions, and transition dynamics can be fully specified in code. Many real-world decisions do not fit this assumption: the decision-relevant state is large and difficult to define, while action consequences are unknown, complex, and context-dependent. We introduce $\textbf{LLM-as-Environment}$, a framework for evaluating open-ended real-world decision problems with structured actions and hard-to-formalize dynamics. At each step, a language model predicts context-dependent consequences of the agent's action, while deterministic code enforces exact mechanics, constraints, and bookkeeping. We instantiate the framework in three domains with $\textit{FuzzyDecisionBench}$: undergraduate career planning, startup strategy, and NFL head coaching. Across $11$ frontier models, the best agents attain $52.2\%$ normalized utility in Career Planning and $24.8\%$ in NFL Head Coaching relative to lower and privileged planning references. We treat simulator fidelity as an empirical question and validate each environment against programmatic simulators, historical trajectories, or human judgments. Replacing YC-Bench's entire Python transition engine with an LLM preserves agent rankings across $8$ models (Spearman $\rho=0.857$); held-out Career Planning transitions achieve $89.0$--$94.1\%$ direction correctness, and simulated NFL tenure title counts correlate with historical outcomes at $\rho=0.825$. More broadly, we view LLM-as-Environment as a promising paradigm for evaluating complex real-world decision making.
Chat is not available.
Successful Page Load