Now or Later: Scheduling Agent Inference Across Real-Time and Batch to Meet SLOs at Lower Cost
Abstract
An agent inference can execute immediately through a real-time streaming API or be deferred to a batch API that cloud providers price at approximately half the real- time rate. Existing agent interfaces generally bind a request to one mode before execution. The runtime therefore cannot revise the mode as conditions change. We identify three requirements for per-inference mode scheduling: (i) the agent definition must abstract the inference mode; (ii) agent execution must be decou- pled from the client-facing API; and (iii) a scheduler must adapt the mode dy- namically to latency and cost objectives. We present CONDUCTOR, a design that satisfies these requirements, and evaluate its scheduler in simulation using param- eterized per-call latency and cost profiles. Given an SLO, the scheduler selects an operating point on the cost–attainment frontier. With exact step-count estimates, an SLO-optimized policy meets every simulated deadline at approximately 21% lower cost than all-immediate execution (cost 0.79; deadline-feasible oracle cost 0.64). When a submitted batch call runs long, the policy launches a duplicate real- time stream at the latest safe time; both requests are charged. The resulting rescue overhead is tunable, decreasing from 13.7% at the cost-minimizing p50 setting to 1.1% at p90. A no-rescue variant trades SLO attainment for lower cost. The task interface supplies an end-to-end SLO, and the runtime selects the execution mode.