Can LLMs Serve as Synthetic Respondents for Consumer Demand Estimation?
Abstract
Large language models (LLMs) are increasingly proposed as "synthetic respondents" for stated-preference research, offering a scalable, low-cost alternative to surveys. This paper asks whether LLM-generated choice data can recover consumer preference parameters and willingness to pay (WTP) for pricing, product design, and other consumer-demand decisions. Using data from a national online discrete choice experiment with 18,144 choice observations, we compare three prompting strategies that provide different amounts of respondent information: pure simulation using demographics and prior purchase prices, few-shot prompting with full choice histories that include both chosen and rejected alternatives, and few-shot prompting that includes only chosen alternatives. First, demographics alone fail to recover economically meaningful preferences. Pure simulation reverses the signs of all product-origin coefficients and substantially overstates the importance of customer reviews. Second, behavioral history improves preference recovery, but both chosen and rejected alternatives matter. Full-set prompting substantially reduces mean absolute percentage error (MAPE) for preference coefficients and WTP, while removing rejected alternatives roughly doubles MAPE. Third, predictive accuracy and MAPE are insufficient to evaluate LLM-generated preferences. Even the best-performing strategy, few-shot full-set prompting, reverses the human ranking of attribute importance, placing customer reviews above product origin, whereas human respondents rank origin first among non-price attributes. Thus, LLM-generated choice data can produce misleading signals for business decisions even when prediction and estimation errors appear acceptable. These findings underscore that LLM-generated responses are not yet a direct substitute for human preference data, but incorporating choice histories that provide trade-off and counterfactual information narrows the gap.