GenAI Evaluation Results are Largely Artifacts of Evaluation Design Choices: Evidence from Audit Studies of Resume Screening
Nicole Meister ⋅ Hannah Cha ⋅ Mohammed Alsobay ⋅ Alexandra Chouldechova ⋅ Alex Dow ⋅ Hanna Wallach ⋅ Jennifer Wortman Vaughan ⋅ Carlos Guestrin ⋅ Tatsunori Hashimoto ⋅ Solon Barocas
Abstract
Generative AI is designed to support many tasks, many ways of framing those tasks, and seemingly endless variation in task inputs. But the same flexibility that makes these systems powerful also makes them difficult to evaluate reliably. This flexibility exacerbates the numerous, but often underrecognized and consequential, \emph{researcher degrees of freedom} where even seemingly minor experiment design decisions can substantially affect findings. We demonstrate this problem through an analysis of 36 papers on gender and racial bias in generative AI-assisted resume screening. Although these studies address the same broad research question, they reach strikingly divergent conclusions. To explain this divergence, we adapt recent work on integrative experiment design from the social sciences, and decompose the dimensions along which the studies vary. We then run 1,134 counterfactual experiments by combinatorially varying the experimental modules (names, job-resume pairings, and models) across three baseline experimental frameworks. We find the results prove remarkably brittle: a different design choice flips the conclusion of the reproduced experiment around 22\% of the time. For example, seemingly inconsequential design choices, such as the phrasing of a job description, can reverse the observed conclusion. However, this instability is not always purely stochastic. In some settings, we can build a model that predicts the result of unseen counterfactual experiments with a mean held-out predictive $R^2$ of 0.807. This suggests that despite sensitivities to difference in evaluation design, there exists a learnable structure.
Chat is not available.
Successful Page Load