How to Train a Surgeon? Benchmarking Generalist Agents in Surgical Scene Understanding
Abstract
Evaluating general-purpose multimodal large language models (MLLMs) for surgical scene understanding remains limited by scarce surgical data and fragmented evaluation protocols. We address these challenges through a two-stage evaluation framework for surgical scene understanding. In Stage I, we introduce SURGKNOWBENCH, a diagnostic benchmark that harmonizes public surgical datasets across procedures and tasks and evaluates proprietary and open-weight MLLMs to map the boundary of pretrained surgical competence. SURGKNOWBENCH reveals that current models retain partial visual priors but remain weak on procedural reasoning and workflow assessment. Building on this failure frontier, and motivated by the scaffolded nature of surgical training, we introduce SURGAGENTGYM, a controlled environment where general-purpose MLLM agents can use bounded knowledge search over SURGWIKI and general-purpose visual grounding for surgical tasks. SURGAGENTGYM evaluates whether generalist agents can improve surgical understanding at inference time through tool-mediated evidence use, without additional surgical training. Our results show that agentic inference narrows the gap mainly for perception-heavy tasks, while fine-grained localization and expert-style assessment remain persistent bottlenecks. This opens a path toward future studies of agentic surgical understanding. We have released the codebase and artifacts.