From Expert Knowledge to Optimization Modeling: Prototype-Based Data Synthesis and Logical Reinforcement Learning
Abstract
Existing approaches to optimization modeling using large language models (LLMs) treat tasks through from-scratch construction, generating variables, objectives, and constraints without leveraging reusable structural knowledge. However, in real-world scenarios, expert modelers typically begin by identifying the canonical type or core structure of a problem, followed by iterative refinements based on specific task requirements. To incorporate this expert incremental logic, we propose a novel two-stage framework. In the first stage, we conduct Logic-Anchored Data Synthesis starting from 17 canonical prototypes, preserving critical bottleneck constraints and generating synthetic problem–formulation pairs. These pairs, together with the OR-Instruct dataset, are used for supervised fine-tuning (SFT) to initialize a policy. Separately, for each generated problem, we produce multiple modeling trajectories that follow a prototype-grounded order with variables preceding expressions. These trajectories are ranked according to criteria of accuracy and efficiency, and subsequently used to train a Logical Reward Model (LRM). In the second stage, we introduce Logic-Test-Time Group Relative Policy Optimization (Logic-TGRPO), which leverages the LRM during test-time reinforcement learning to reward accurate prototype identification and disciplined structural adherence while penalizing illogical patterns. Evaluated across seven benchmarks, our 8B parameter model achieves an average accuracy of 79.9%, outperforming comparably scaled baselines and competing effectively with heavyweight multi-agent systems with far fewer parameters. Strong performance on out-of-distribution tasks confirms that the model exhibits genuine incremental reasoning rather than mere memorization. Ablation studies further validate that both stages in our framework are essential.