EchoGemma: A Modular Vision-Language Model for Echocardiography Report Generation and Visual Question Answering
Abstract
Echocardiography is the most common form of cardiac imaging, but its interpretation is subjective and its reporting is verbose and repetitive. Medical vision-language models are designed around single images and perform poorly on cardiac ultrasound, where a study comprises dozens of videos across multiple views. We present EchoGemma, a vision-language model that adapts MedGemma to echocardiography by first writing the findings as a clinical report and then answering questions from that report. A domain-specific video encoder reduces an entire study to view-aware embeddings; these are projected into a LoRA-fine-tuned medical language model, which generates a structured clinical report. Because the report is written in clinical language, answers can be checked against the stated findings, and errors can be traced to either the report or the answering step. We evaluate report generation on an in-house study-report dataset and question answering on MIMICEchoQA and on a new cardiologist-reviewed benchmark derived from MIMIC-IV-ECHO. Across 11 interpretation tasks, EchoGemma increases mean accuracy from 46.7\% to 93.5\% and reduces the mean absolute error on left ventricular ejection fraction from 16.0\% to 5.17\% relative to MedGemma. Additionally, supplying the generated report as context improves question answering accuracy from 56.4\% to 59.2\% and from 57.8\% to 64.6\% on the two benchmarks. Our results indicate that grounding a medical language model in study-level video improves echo interpretation, and that an explicit clinical report is a useful and interpretable intermediate step for visual question answering.