EndoSCOP-V: A Multi-Turn Video Understanding Evaluation Framework for Multimodal Models in Endoscopy Reporting and Clinical Reasoning
Abstract
Real-world endoscopic diagnosis unfolds across native video, depending on continuous traversal of gastrointestinal organs where a single positive finding triggers a multi-minute reasoning chain over location, morphology, severity, and intervention. Existing endoscopy multimodal benchmarks evaluate only static images or single-disease videos that fail to capture the temporal and clinical reasoning demands of real-world endoscopy reporting. We introduce EndoSCOP-V, a video-grounded multi-turn evaluation framework for long-horizon clinical reasoning in endoscopy. The benchmark contains 1,307 real-world endoscopy and colonoscopy videos (63.2 h; 97% 1080p; mean ~174 s) — the time-scale at which endoscopists actually reason about a finding. The corpus is paired with 20,275 question–answer turns deterministically instantiated from clinician-filled structured endoscopy-reporting EHRs, avoiding synthetic or LLM-rephrased per-case annotations. Questions are derived from 1,033 clinician-reviewed templates spanning a clinically aligned 9-task × 30-sub-task reporting taxonomy across 54 World Endoscopy Organization (WEO) MST 3.0 diseases at ~18 templates per disease, substantially deeper per-disease examination than prior endoscopy benchmarks. To mirror real-world case mix, 22% of procedures carry co-existing pathologies, up to six per case. To model the hierarchical and conditional nature of clinical reporting, EndoSCOP-V introduces a Markovian context-pruned oracle-forcing protocol that evaluates multi-turn dialogues turn-by-turn. By injecting the gold answer before each subsequent turn, the protocol isolates per-turn reasoning from cascading dialogue errors and enables fine-grained analysis of clinical reasoning capabilities. Clinician-marked key-frame anchors further separate temporal-localization from frame-level comprehension failures. A panel of 15 board-certified endoscopists anchors a 78% strict-accuracy baseline on a 1,500-question audit. Across 14 open-weight LMMs, the best 27B-class model reaches 59.8% strict accuracy, falling to 53.4% once format-trivial disease-rejection items are excluded. On a between-subjects spot-check, the best of 3 frontier proprietary models trails the clinician panel by 23 pp on strict accuracy. Sub-task analysis reveals consistent gaps in quantitative measurement, multi-variable grading, and intervention reasoning on long-horizon video, providing a clinically grounded benchmark for evaluating multimodal reasoning in endoscopy.