FirmVLA: Sequence-SLO Serving for Heterogeneous VLA Fleets
Ahnaf Adib ⋅ Sabbir Ahmed ⋅ Latifur R Khan
Abstract
Vision-language-action (VLA) policies differ in inference time, batch capacity, action-chunk length, and refill rate, so robots that share a GPU need service on different schedules. We present FirmVLA, a GPU serving system for mixed VLA policies. Each robot specifies an $(m,k,r)$ sequence service-level objective (SLO): serve at least $m$ of every $k$ refills and miss no more than $r$ in a row. FirmVLA builds an exact repeating admission schedule and preserves its reservations while serving. It records zero SLO violations on six fleets with two to four policies and 14–122 admitted robots, and in 14 trials across start times and loads from 50% to 100%. In those trials, baseline schedulers complete the same total work and still violate SLOs at high load. The same split holds with two policies competing for GPU slots and batch places. FirmVLA stays at zero violations, while DBP, EDF, round robin, and Armory Lookahead violate SLOs at high load. Across the same 3,584 task episodes on four policies, robots that keep leftover actions through missed refills succeed 79.8% of the time, against 41.2% when they discard them.
Chat is not available.
Successful Page Load