Compute Allocation Under Model-Provider Competition
Abstract
Competition among LLM providers hinges not only on model accuracy, but also on the latency of serving user requests. However, user demand leads to congestion which increases latency, creating a feedback loop between provider decisions and user choices. In this work, we study a stylized game between two providers who strategically balance allocating fixed compute between training better models and reducing latency. A population of users then chooses between the providers. Our equilibrium analysis reveals that a lower-budget provider can still attract users by taking advantage of the congestion faced by the dominant provider. Relative to a competitive baseline where providers have equal compute resources, compute asymmetry leads the higher-resource provider to divert a greater fraction of their compute from improving accuracy to reducing latency. Moreover, this compute asymmetry strictly harms users. Altogether, our results illustrate how congestion distorts competition between model-providers, leading to subtle implications for market competitiveness, model accuracy, and user utility.