Back in Style: A Sociolinguistic Approach to Authoring and Measuring Persona Fidelity in User Simulation
Abstract
User simulators increasingly serve as the measurement instrument behind agentic evaluation, yet the fidelity of simulated users in comparison to real human users is generally low and typically assessed by costly, subjective LLM judges. In this pilot study, we ask whether fidelity can instead be measured deterministically by treating a user persona sociolinguistically: as a social type that emerges from observable linguistic style, rather than one predicted by labels or descriptions a model must extrapolate into behaviour. Authoring personas as concrete stylistic rates lets us transfer two established, model-free instruments - authorship-verification stylometry and lexicon-based content analysis - as persona fidelity diagnostics. We A/B-test this schema against a flat descriptive baseline across 5 task-oriented customer-service agents. Results show that the sociolinguistic schema improves how faithfully the simulator renders persona style in terms of overall feature rates, with a caveat that persona style fidelity does not equal persona authenticity. We argue that the metrics prove most useful as an error-attribution analysis, localizing where fidelity breaks down, as a first step towards interventions that would move user simulations towards more faithful renderings of diverse and variable linguistic outputs.