MAPS: Margin-Aware Priors and Verifier-Guided Search for Embodied Planning
Abstract
Large language models (LLMs) have shown strong reasoning capabilities in embodied planning tasks. However, they often fail to achieve good results on long-horizon and multi-robot tasks under strict physical constraints, and generate smooth but infeasible trajectories. In real-world robotic hardware deployment, systems also face two fundamental challenges: First, the fine-tuning method for open-source models, Direct Preference Optimization (DPO), typically only makes a binary true or false judgment on the planning results, without distinguishing the severity of the planning error. As a result, they fail to impose effective penalties for catastrophic violations of physical constraints. Second, methods that search over multiple candidate plans during inference do not filter out invalid plans in advance, often leading to combinatorial explosion and high inference latency, which makes them difficult to apply to real-time robotic operations. To address these limitations, we propose a complete training-to-inference architecture consisting of two modules. First, we introduce Margin-Aware Direct Preference Optimization (mDPO) during training, which distinguishes the severity of errors in negative samples, enabling the fine-tuned model to avoid generating severe errors that would render a task irreparable. Second, we deploy Verifier-Guided Search (VGS) during inference, using a symbolic verifier to hard-prune invalid search paths, thereby repairing infeasible task plans and reducing inference latency. In addition, latency analysis shows that our approach reduces the average end-to-end runtime from 12.814s to 10.969s compared with retry-based repair, while maintaining higher final success, making it more suitable for time-sensitive robotic deployment.