SPRING: Solver-guided Process Rewards for Novel Logical Reasoning Steps Generation
Abstract
Logical reasoning remains a major challenge for large language models (LLMs), particularly on structured problems that require precise constraint tracking, consistency preservation, and multi-step deduction. This challenge is especially acute for small-scale LLMs, which are more prone to producing inconsistent, redundant, or brittle reasoning trajectories. Existing approaches for improving logical reasoning largely optimize for final-answer correctness, providing only weak supervision over the intermediate reasoning process. In this work, we propose SPRING, a solver-guided reinforcement learning framework for logical reasoning that uses an SMT solver as a training-time verifier of intermediate reasoning steps to provide process-level supervision. SPRING introduces the notion of a novel reasoning step, namely, a step that is logically valid, consistent with the evolving reasoning state, and not already implied by previously accepted non-contradictory deductions. Based on this solver-based assessment, we design process rewards that encourage novel inferential progress while penalizing contradictory and uninformative reasoning steps. Evaluation results on two logical reasoning benchmarks, ZebraLogic and AR-LSAT, show that SPRING consistently outperforms baseline LLMs and outcome-only reward baselines. On ZebraLogic, SPRING improves puzzle accuracy by up to 49.71 and 15.43 points over the base LLM and outcome-only reward baseline, respectively. On AR-LSAT, SPRING improves the overall average score by up to 64.93 and 12.14 points over the base LLM and best-performing outcome-only reward baseline, respectively.