Diagnosing Empirical Target Retention in Multi-Turn LLM Scenario Construction for Household Food Waste
Abstract
Detailed case-level accounts of household food waste are scarce at scale, while available empirical evidence is often partial or aggregate. LLMs offer a way to elaborate such evidence into contextualized synthetic scenarios, but it is unclear whether assigned empirical properties remain represented as those scenarios become richer. We test this using 100 cases assigned a food category and primary waste reason, comparing one-shot generation (SINGLE) with progressive construction through the same 31 questions in randomized order (Random MULTI). Random MULTI ended with lower fidelity than SINGLE for both food (65.7% vs. 82.0%) and reason (35.3% vs. 45.0%), and greater marginal divergence (food TVD: .297 vs. .170; reason: .543 vs. .360). Most loss accumulated during elaboration rather than final synthesis, with reason declining more strongly than food (INITIAL--Turn 31: -.531 vs. -.207); losses were also positively associated across cases (rho=.414). These results show that externally assigning empirical properties does not ensure their retention as generated scenarios become more contextualized and sequentially elaborated.