When Can We Trust LLM Personas? Stress-Testing Behavioral Validity Beyond a Single Scenario
Jiwoo Choi ⋅ In Ik Lee ⋅ Woo Sung Seo ⋅ Minho Lee ⋅ Yun Young Choi
Abstract
Large language models (LLMs) are increasingly used to simulate persona-conditioned human behavior. However, whether behavioral fidelity observed in a single scenario transfers to new environments remains underexplored. We study this question in cold-start e-commerce advertising, using real Google Ads outcomes as ground truth for segment-level click behavior. Starting from a persona simulator developed on one product, we freeze its configuration and evaluate it under temporal and product shifts. Under a four-month temporal shift on the same product, the simulator remains relatively robust, achieving an MAE of 31.2—53\% lower than the best tested non-LLM baseline. In contrast, it fails to transfer reliably across two held-out products, producing errors approximately 16$\times$ and 1.8$\times$ those of their respective best non-LLM baselines. Additional stress tests show that reversing the injected behavioral cue substantially alters population ordering, prompt paraphrases can yield large performance variation, and calibration can fail across LLMs. These results show that strong in-scenario accuracy does not by itself establish behavioral validity, motivating evaluation across held-out scenarios, behavioral-prior interventions, prompts, and models.
Chat is not available.
Successful Page Load