Doctor-in-the-Loop Generative AI for Real-World Physician-to-Physician Teleconsultation
Abstract
Specialist maldistribution remains a major constraint on sustainable care, and teleconsultation can broaden access only within the limits of available specialist time. We implemented and evaluated a doctor-in-the-loop generative AI system using consultations from a physician-to-physician teleconsultation service piloted in real-world clinical practice at 20 institutions. We evaluated feasibility as a mechanism for delivering specialist knowledge into practice and the complementary roles of generative AI and specialists in reply quality. Of 37 consultations, 35 (94.6%) were resolved without an in-person referral, and the median time to first response was 9.2 hours. Reviewing the AI-generated structured record took a specialist a mean of 87.0 seconds per consultation, corresponding to an estimated 77.4% reduction in documentation time. In resident-initiated consultations, the model-generated reply was more comprehensive than the specialist’s own (4.17 versus 3.24) but had lower clinical applicability (3.31 versus 4.26). Specialist editing retained high comprehensiveness (4.29) while improving clinical applicability (4.02). These findings suggest that doctor-in-the-loop generative AI may extend specialist capacity by supporting routine information processing while preserving specialist involvement in clinical judgment.