A Deployed Emotional-Support Dialogue System with Auditable Interaction Evidence
Abstract
Large language models make supportive dialogue systems easier to prototype, but real deployment also requires continuity, reliable operation, and evidence for claims about human-agent interaction. We report on our deployed public-facing supportive dialogue system for everyday emotional support and conduct a validity audit of two generations of evidence from its real-world use. The system combines cloud-hosted LLM agents, expert-in-the-loop prompt iteration, stage-aware support, and backend services for routing, message storage, user-profile summaries, and reporting. Using de-identified data, we compare a six-month general-public deployment comprising 31,823 message-level records from 1,342 users with a prospectively instrumented longitudinal field experiment in which 100 participants produced 8,871 completed turns. The general-public records show deployment scale and operational behavior but cannot relate individual system decisions to users' expressed needs. The prospective records made all 867 observed stage transitions reconstructable by source and direction. We linked all 622 system-involved transitions from the study's log-analysis subset to human assessment. Although the corresponding event records were internally consistent, 547 transitions (87.9\%) were judged appropriate, 26 (4.2\%) inappropriate, and 49 (7.9\%) unnecessary. This finding shows that a system decision can be fully traceable without being appropriate. Based on this comparison, we organize deployment evidence into four levels: implementation, operational, interaction-process, and outcome. We also propose a per-turn event-record schema that identifies the system records and human judgments needed at each level. Together, our results show how event-level records and linked human assessment can support stronger evaluation of deployed human--agent systems without treating observability as evidence of interaction quality or user outcomes.