Scripto: An On-Premises Clinical Speech Platform for Language-Adapted Dictation and User-Defined Documentation Workflows
Abstract
Clinical documentation remains a major source of physician workload, yet hospital deployment of speech and ambient-scribing systems is constrained by local language, institution-specific terminology, data-governance requirements, and fragmented clinical software. We present Scripto, a fully on-premises clinical speech-to-documentation platform that unifies low-latency dictation, local ASR adaptation, and clinician-configurable, source-grounded generative workflows. Scripto is currently in hospital-wide pilot deployment at a large medical center and processes approximately 24 hours of dictation per week. Its application-independent client inserts recognized text at the active cursor in arbitrary clinical software. With user consent, locally collected recordings are incorporated into periodic model updates for institution-specific language. Beyond dictation, a shareable workspace combines live speech, uploaded audio, transcripts, documents, and natural-language instructions into reusable templates for long-form transcription, speaker diarization, and structured note generation. Generated content is grounded in quoted transcript passages linked to corresponding audio segments and can include citations to retrieved clinical guidelines, facilitating verification. We describe the system architecture and adaptation pipeline, report speech-recognition results, and demonstrate two end-to-end workflows: generating SOAP notes from clinician–patient interviews and ambient procedural documentation with retrieval-augmented generation. Scripto illustrates a practical path toward institution-controlled, adaptable, and verifiable clinical speech AI.