Too Agreeable to Teach: Multi-Resolution Evaluation of Simulated Learners in AI-Facilitated Training
Abstract
Community mental-health training in rural India relies on scarce facilitators running group sessions where learners surface uncertainties and misconceptions that are resolved through dialogue. We study a smartphone AI facilitator leading simulated peer learners alongside a focal learner. This is a multi-party rather than single-user simulation: one simulator's contribution changes the evidence available to every subsequent participant. Grounding peer states in human group training organized into dialogic episodes, we evaluate at three resolutions: peer invocation, interaction arc, and multi-session trajectory. The simulated classroom reaches curriculum milestones faster than the human training, but through simulator-assisted task advancement rather than learner progress: the facilitator exposes the target, cooperative peers return aligned evidence, and that evidence is ratified as learning; at times the facilitator doesn't use peers at all. Judges certified this pattern, accepting peer invocations across the simulation trajectory. Replacing holistic judgment with transcript-anchored atomic questions most of audited failure targets. Evidence provenance and state-transition validity are first-class evaluation properties.