AURA: Amortized, User-Conditioned Reranking Adapters for Personalized LLM Agents
Abstract
Personalizing LLM agents typically requires injecting a user profile into the prompt at every request, increasing context length and inference cost. We investigate an alternative paradigm: amortizing personalization into model weights. We present AURA (Amortized, User-Conditioned Reranking Adapters), a hyper-network that maps an account-profile embedding to a low-rank adapter for a frozen instruction-tuned LLM. The adapter is generated once per account and can then be cached, so subsequent ranking requests contain only the query and candidate products. We study AURA on personalized product reranking using naturally occurring queries and purchase outcomes from a cloud marketplace. Across six instruction-tuned models spanning three model families and 1.3B–8B parameters, AURA achieves the highest Hit@1 point estimate on all six models. Relative to query-only ranking,AURA improves Hit@1 by 7–30 points across models. Compared with in-context profile injection, AURA provides a clearer accuracy advantage on some models, while matching its accuracy on others at substantially lower latency. Profile injection increases per-request latency by 43–265%, whereas AURA adds only0.2–2.8%; because the generated adapter depends only on the account profile, it can be precomputed and cached. We further examine how account information is represented in the generated adapters. Semantic profile embeddings consistently outperform sparse categorical representations. Interestingly, cross-entropy training produces highly similar adapters across accounts, with pairwise cosine similarity of 0.986–0.998 in our study on DeepSeek-Coder-1.3B. Despite this similarity, the adapters can still induce different rankings across accounts, indicating that parameter-space similarity does not imply identical behavior. Adding a cosine-diversity objective substantially increases adapter diversity and produces more account-specific rankings while preserving mean Hit@1 in that experiment. Together, these results suggest that weight-space personalization is a promising alternative to repeated profile injection for latency-sensitive LLM agents, while raising an important question about how account-specific information is distributed within generated adapters.