Med-Agentic: Distilling Agentic Medical Reasoning with Internalized Meta-Capabilities
Abstract
Large language models (LLMs) in expert domains face a trade-off between \emph{agentic tool use}, which externalizes expertise behind brittle multi-turn runtimes, and \emph{end-to-end specialization}, which internalizes expertise only implicitly at task-level granularity. We propose \emph{internalized meta-capabilities}: atomic, parametrically internalized, and composable reasoning units that an LLM can select and combine within a single chain-of-thought (CoT). We instantiate this idea as \textbf{Med-Agentic} for medical visual diagnosis, where the model emits a capability-tagged structured CoT in one forward pass, turning capability invocation into tagged token generation rather than external tool calls. For training, \textbf{Med-OPD} performs multi-teacher on-policy distillation with \emph{token-level dual-factor teacher routing} conditioned on both the image and the active capability tag, strictly generalizing per-prompt teacher routing. We show that this routing gives an unbiased estimator of the same expected teacher-KL gradient. Across public medical imaging datasets spanning nine anatomies and eight modalities, Med-Agentic consistently improves in-distribution and out-of-distribution diagnostic accuracy over the strong competitors.