MedPsy: State‑of‑the‑Art Small Medical Language Models for Efficient Edge Deployment
Abstract
Medical language models promise expert clinical support, but the strongest open source LLMs remain too large for private, edge devices deployment, while compact medical models still trail substantially on knowledge-intensive and real-world healthcare tasks. We present MedPsy, a family of text-only 1.7B and 4B medical language models designed for edge deployment. Our recipe combines a synthetic medical data pipeline over biology, medicine, and health seeds with chain-of-thought targets from a 235B medical-focused LLM teacher; a four-stage post-training curriculum of two SFT stages followed by two RL stages with hard-sample mining; and a mobile-oriented quantization study over several GGUF variants per model. On seven closed-ended medical benchmarks, MedPsy-4B surpasses MedGemma-1.5-4B by +19.34 points and matches MedGemma-27B while being 6.75x smaller; MedPsy-1.7B outperforms MedGemma-1.5-4B by +11.42 points despite being less than half its size. On HealthBench-Hard, MedPsy-4B surpasses MedGemma-27B by +15.33 and MedPsy-1.7B surpasses it by +11.66 points despite being 16x smaller. Beyond accuracy, MedPsy reduces average response length by 1.7x (1.7B) and 3.2x (4B) versus its Qwen3 backbones, and 4-bit quantization retains accuracy within 1 point of BF16 while reducing disk footprint by ~69%, enabling practical deployment on resource-constrained edge devices.