Medical LLMs as Medical World Models: Unified Policy-Dynamics Learning with Test-Time Search
Yucheng Zhou ⋅ Peng Luo ⋅ Jianbing Shen
Abstract
Clinical decision-making is intrinsically sequential: a clinician must propose a diagnosis and treatment plan from the pre-treatment state and anticipate the post-treatment outcome in order to revise the plan. Existing medical world models address this loop only partially, through image-to-image dynamics with separate diagnostic and scoring modules, text-only trajectory models without multimodal input, or reflection agents without a world model. We propose \textbf{MedWM}, a unified multimodal LLM in which policy and dynamics are two modes of a single autoregressive backbone with fully shared parameters; the dynamics mode emits a text-only \emph{structured} post-treatment state, free-text clinical description plus standardized key-value outcomes (pCR, RECIST, labs, survival), that admits objective evaluation without perceptual metrics. We post-train this backbone with joint policy-dynamics supervision and dynamics-grounded reinforcement learning, then wrap it in an inference-time agent harness that combines breadth-depth ($N \times K$) search using the model's own dynamics as an internal verifier with a self-evolving case memory. On MIMIC-IV, BreastDCEDL-ISPY2 and HCC-TACE-SEG, MedWM consistently improves over same-backbone controls, matches or surpasses a specialized three-module image-based world model, and outperforms reflection-only, multi-agent, and larger closed-source LLM baselines; a blinded clinician study corroborates the quantitative gains.
Chat is not available.
Successful Page Load