The Value of Personalization in Large Language Model Responses
Abstract
We study how personalized LLM routing increases end-user utility. Using Arena human-preference data, we estimate a random-coefficients demand model and use it to simulate the effects of a router that scores responses from ten models and picks the one the user, prompt, or user-prompt pair is predicted to prefer. We find that a prompt only router raises predicted mean per-user utility by 1.16 logits over an always-GPT-4o baseline, attaining a 67.8\% model-implied head-to-head win rate. Allowing the router to personalize to each user's own tastes, rather than serving everyone the average user's pick, adds a further 0.83 logits, for a total gain of 1.98 logits and an 81.2\% win rate. The allocation of responses includes all ten models we consider, which makes routing worthwhile and highlights that measured intelligence is not the only relevant model characteristic. Next, we investigate the feasibility of learning individual preferences using repeat interactions. We find that after forty feedback rounds, we can close 68% of the gap to the optimal router. Lastly, we consider the efficacy of verbalizing estimated tastes. This approach gains 0.92 logits for the intended user (64.9% predicted preference), which falls short of routing.