RADIUS: Ranking, Distribution, and Significance - A Comprehensive Alignment Suite for Survey Simulation
Abstract
Simulating survey responses with LLMs offers a scalable alternative to costly human data collection, yet evaluation methods remain ad hoc and inconsistent. Current metrics, typically borrowed from adjacent domains, focus on accuracy or distributional similarity while overlooking ranking alignment—whether simulations preserve the relative preference orderings observed in human populations. A simulator may achieve low distributional divergence yet fail to identify the option humans most prefer, a critical gap for decision-making applications. We propose RADIUS, a two-dimensional alignment suite that evaluates survey simulation along complementary axes: 1) RAnking alignment and 2) DIstribUtion alignment, accompanied by statistical Significance testing. Experiments across multiple social surveys and persona sources demonstrate that RADIUS exposes limitations of single-metric evaluation and offers stronger discriminative power. On an independently reproduced simulator, we further show that distributional improvement often coincides with misidentifying the option humans most prefer—a decision-relevant error RADIUS surfaces but distributional metrics miss. We open-source RADIUS to support reproducible and comparable assessment of survey simulators.