Multi-Turn RL Makes Small Language Model Competitive for Optimization Modeling
Abstract
Mathematical optimization drives decisions in supply chains, logistics, energy, and scheduling, but translating natural-language problems into solver-executable formulations remains a bottleneck. While frontier models show promise for this task, many deployments require small language models (SLMs) that run locally under privacy, latency, and cost constraints. Existing SLMs are limited by noisy data, ambiguous benchmarks, and one-shot inference that ignores the iterative, feedback-driven nature of optimization modeling. We introduce \textsc{OptiMind}, a framework for improving SLMs through formulation-centric data cleaning, class-conditional priors, and multi-turn reinforcement learning with solver feedback. \textsc{OptiMind} repairs ambiguous problems, regenerates solution trajectories, audits benchmark labels, constructs accepted objective-value sets, and distills class-level hints from recurring formulation errors. It then trains models to formulate, execute, and revise GurobiPy programs using solver-verified correctness as reward. On expert-cleaned IndustryOR, Mamo-Complex, and OptMATH benchmarks, a 20B \textsc{OptiMind-RL} model outperforms strong open-source baselines and reaches frontier-model performance under the same protocol. Our results suggest that optimization formulation is better treated as an interactive solver-grounded decision process than as one-shot code synthesis. We release our inference framework, cleaned benchmarks, and sample error analyses at \url{https://anonymous.4open.science/r/OptiMind-NeurIPS-1DB020}, with full data and model to follow.