Benchmarking Continual Learning in Orchestration Agents
Abstract
Enterprise LLM agents complete requests by orchestrating across services that different teams own. Many of those requests repeat a similar workflow with different inputs, so an agent that can continually learn from its past requests can complete tasks faster and more reliably. At the same time, service owners independently change the organizational rules governing these workflows, so agents must recognize when what they learned from earlier requests no longer applies and adapt. Recent benchmarks study continual learning in LLM agents or adaptation to evolving tools. Yet they do not test a distinct, practically important challenge: learning to orchestrate repeated cross-service workflows while detecting changes in the organizational policies that govern them. To study this challenge, we introduce Continual Orchestration Bench, a benchmark which presents an agent with an ordered sequence of similar requests over simulated enterprise services, grades each request on the final state of those services, and partway through changes the written rules, which the agent must discover through its tools. The benchmark covers eight enterprise workflow families, 128 requests, and 226 native tools. Our baseline evaluation covers four of these families, yielding a frozen 64-request stream. We compare the same model operating without memory, using in-context learning over its past requests, and using external memory systems. To establish that information from earlier requests can improve later performance, we include a privileged control that retrieves the most efficient relevant prior success. Both in-context learning and external memory methods improve mean reward over the model without memory, but both lag behind the control, showing that current systems capture only part of the benefit available from learning across requests. Our results show that common memory methods, which perform well on static retrieval benchmarks, struggle to adapt and recover under rule changes, motivating better continual-learning methods for enterprise orchestration agents.