CEO-Bench: Can Agents Play the Long Game?
Abstract
Language model agents are becoming proficient executors at isolated, short-horizon tasks such as software engineering and customer service. Yet real-world challenges require a combination of sophisticated skills that remain largely untested in agents: (1) navigating long horizons amid uncertainty; (2) acquiring information in noisy environments; (3) adapting to a changing world; (4) orchestrating multiple moving parts toward a coherent goal. We introduce CEO-Bench, which evaluates these capabilities together by simulating a representative real-world task: operating a startup for 500 days. Given diverse and realistic company management tools, business databases, and social media, an agent needs to design pricing strategies, allocate operating budgets, analyze business data, respond to unexpected competitor moves, and more. Our evaluation shows that most state-of-the-art models struggle to succeed in this environment, and only one model (GPT-5.5) finishes the simulation above its $1M starting balance. CEO-Bench reimagines the role of AI in the future, shifting from solving isolated tasks to driving sustained, adaptive progress over time. We open source full code and trajectory.