Persona Generators: Generating Diverse Synthetic Personas for Arbitrary Contexts
Abstract
Simulating human behavior with Large Language Models (LLMs) offers a scalable laboratory for social science and stress-testing AI products. However, limitations remain in matching LLM outputs against the full breadth of human diversity. Current techniques for creating synthetic humans typically optimize for density matching, which often enforces behavioral uniformity and overlooks the rare, consequential outliers. We argue that robust simulation requires shifting the focus toward support coverage: ensuring synthetic populations span the entire landscape of possible human traits, opinions, and preferences. This paper introduces Persona Generators, functions that can produce diverse synthetic populations tailored to arbitrary contexts. We apply an iterative improvement loop based on AlphaEvolve, using LLMs as mutation operators to evolve our Persona Generator code rather than the personas themselves. The optimization process produces lightweight functions that can automatically expand small descriptions into populations of diverse synthetic personas that are maximizing coverage along relevant diversity axes. We demonstrate that evolved generators substantially outperform existing baselines across six diversity metrics on held-out contexts. Furthermore, by successfully mitigating standard LLM mode collapse, our generated populations actually capture real human trait distributions better than baselines specifically designed to match human statistics.