MedSkill: In-Context Skill Learning for BioMedical Tool-Use Agents
Abstract
Large language models are increasingly used as agents in healthcare related tasks. The heterogeneity of such biomedical tasks usually requires adapting a separate policy to each task. However, there is often similarity between procedures used across tasks, motivating mechanisms that learn skills by gathering evidence and testing alternatives. We propose MedSkill an in-context training-free skill-learning framework that iteratively designs and refines a shared bank of natural-language procedures and composes them to solve subsequent tasks more efficiently. Skills are optimized for reuse across a distribution of tasks, while the controller learns task- and state-dependent composition, enabling specialization without training a separate end-to-end model per task. Its actor and evolver remain frozen, and its prompt-based interface requires only black-box model access. We instantiate the framework on Athena-style healthcare evaluations, including biomedical knowledge, and clinical decision tasks, and compare skill-composed policies with direct deployment of the underlying language models.