Why Personalized LLM Agents Fail at Implicit Preference Updates
Abstract
Personalized language model agents must dynamically adapt to evolving user preferences, which are often implicitly necessitated by life events rather than explicitly declared. To evaluate this capability, we introduce a diagnostic framework that isolates implicit event-to-preference inference from basic memory retrieval. Across nine models and 27,720 items, we reveal a systemic failure in implicit memory supersession: while models successfully follow explicit preference updates (near 100\% accuracy), they fail catastrophically at inferring those same updates from causal events, performing significantly below chance (2.2\% to 13.3\%). We demonstrate this is neither a long-context retrieval failure nor solvable by increased test-time compute. Blind-scored evaluation of reasoning traces reveals profound chain-of-thought unfaithfulness; models that successfully derive the event's consequence still fail to update their final decision. Instead, interventions reveal the failure stems from rigid anchoring to explicit priors. Artificially weakening the initially stated preference improves implicit inference but conversely causes models to improperly override explicit constraints. Contrastingly, human annotators easily resolve these implicit updates (92.9\% accuracy). By demonstrating that current models cannot implicitly supersede memory, this work exposes a critical bottleneck in robust agentic personalization. We will release the dataset, model outputs, and analysis code.