Wrong Together: Shared Bias in User Simulation
Saber Zerhoudi
Abstract
A user simulator stands in for the participants of a study, so the study's result can be estimated without running it on humans. But whether a given simulator was right is known only once the human data it was meant to replace are collected, and by then it is too late to act on. We ask whether the disagreement among several simulators can warn of those errors before the human data exist, and find that the answer depends on what the study measures. For a single group's response, much of the error is a bias the simulators share, which their agreement hides, so disagreement predicts that error less reliably. For the difference between two groups, that shared bias cancels, and disagreement predicts the remaining error well. We trace this to the shared bias itself: estimating it on separate data and subtracting it corrects the single-group estimates, while leaving the disagreement signal unchanged. On a public survey of opinions across dozens of countries, run through six simulators, we let the simulators answer the differences between two groups where they agree and send the rest to humans; the answered differences have a $54.8\%$ error rate, down from $63.7\%$ when the simulators answer all of them.
Chat is not available.
Successful Page Load