Market-Based Runtime Resource Allocation for LLM Multi-Agent Systems
Abstract
LLM multi-agent systems rely on division of labor to solve complex tasks, but their execution can be blocked by contention for shared resources such as GPU test environments, CPU execution slots, and exclusive tool sessions. Existing workflows, task allocation methods, and runtime schedulers usually arbitrate such access indirectly through workflow position, dependency structure, queue state, or resource metrics. These criteria do not fully capture the value of satisfying a specific request, because request value is distributed across agents' local execution states and changes with tests, retries, and downstream results. We propose a market-based protocol for access arbitration in LLM multi-agent systems. Each agent converts its current plan and local execution state into a structured, budget-constrained bid that signals relative urgency. For asynchronous requests, an online clearing method based on optimal stopping decides when to close the clearing window, allocate access rights, settle payments, and track resource occupation until release. Across three contention scenarios, Market reduces makespan by 5.2\%, 2.5\%, and 4.9\% relative to the strongest primary baselines, while improving task score by about two percentage points.