ARES: How Reliable Are LLM User Simulators for Recommender A/B Testing?
Hongyang Su ⋅ Beibei Kong ⋅ Lei Cheng ⋅ Chengxiang Zhuo ⋅ Zang Li ⋅ Chenyun YU
Abstract
LLM-driven user simulators are increasingly adopted for A/B-testing recommender systems, yet the choice of backbone is treated as an interchangeable implementation detail rather than a factor shaping evaluation reliability. We show otherwise: across nine backbones under unified controls, this choice accounts for 149% as much outcome variance as the tested recommender, and swapping backbones alone flips about 20% of pairwise A/B winners. The backbone, in other words, is not a neutral component but the measurement instrument. We turn this observation into a method: ARES (Audit for Reliable Evaluation of Simulators) separates a backbone effect (the instrument's contribution) from a recommender effect (the signal under study), quantifies both at three levels (outcomes, trajectories, and reasoning), stress-tests them against text vs. rendered UI, and endorses only A/B claims a majority of backbones independently support. Applied to nine backbones $\times$ five recommenders on MovieLens-1M, ARES surfaces failures invisible to outcome metrics: within a single model lineage, newer and stronger releases do not necessarily agree more with their predecessors (Kendall's $\tau$ drops from 1.0 to 0.20 across generations); three of eight vision-capable backbones reverse the declared best recommender under UI rendering; and 63000 reasoning traces tie these reversals to interpretable behavioral biases, not stochastic noise. Capability alone, therefore, does not guarantee reliability. While instantiated for recommendation, the backbone-as-instrument view generalizes to any LLM-mediated evaluation where model identity is a free variable, including LLM-as-a-Judge. We release ARES-Bench at https://github.com/neurips2026-ares-authors/ARES-Bench: 21400 behavioral logs, 63000 reasoning traces, a reusable visual sandbox, and an analysis toolkit, making multi-backbone auditing a feasible default for simulator-based evaluation.
Chat is not available.
Successful Page Load