Conversational Psychometrics: Toward Validity-Centered Measurement of Human Epistemic Agency in Human–AI Dialogue
Abstract
Evaluation of human–AI collaboration has concentrated on one side of the team: we benchmark models exhaustively, yet we measure the human's contribution with self-report scales or not at all. Team-level outcomes conflate what the model did with what the human contributed, so we cannot tell whether a deployed system amplifies human judgment or quietly replaces it—the central question of human–AI coevolution. This position paper argues that the missing measurement instrument already exists in the interaction itself: dialogue logs provide naturally occurring, large-scale behavioral traces of partially externalized reasoning, in which questioning, verification, self-correction, and the regulation of delegation become observable. We propose conversational psychometrics: the measurement of a person's opportunity-conditioned epistemic conduct in human–AI dialogue—with person-level epistemic agency treated as a further inference requiring cross-task, cross-model, and AI-absent transfer evidence—built on evidence-centered design and held to the validity standards of educational and psychological measurement rather than the conventions of benchmark leaderboards. We define the target construct as four behavioral dimensions (problem framing, evidence evaluation and omission awareness, revision in response to counter-evidence, and regulation of delegation and responsibility), specify an evidence pipeline with pre-stated reliability and validity thresholds, and identify three methodological challenges that make this setting unlike existing assessment paradigms: partial observability of cognition, decomposition of human ability from human+AI system performance, and metric drift under model updates. A deployment vignette from an AI-assisted grading workflow, where instructors' final judgments diverged systematically from the workflow's initial scores, serves only as motivation, because those scores cannot always be attributed to the AI side; validation of the proposed dialogue-based instrument is charted as the program's first milestone rather than claimed as accomplished. We close with a research agenda and governance principles (non-punitive use, return-to-learner) intended to keep such measurement aligned with the people it measures.