DiagSQL: A Diagnostic Validator for Text-to-SQL with Reward Allocation and Co-occurrence Shaping
Abstract
While large language models have significantly advanced Text-to-SQL generation, it remains largely an open-loop paradigm due to the absence of a reliable validator capable of systematically diagnosing generated queries. Training a dedicated diagnostic validator to provide targeted, reflective feedback is crucial for closing this loop. However, employing standard Reinforcement Learning paradigms, particularly Group Relative Policy Optimization (GRPO), reveals fundamental structural misalignments. Through exploratory analysis, we identify a critical limitation of standard GRPO in this context: a severe credit assignment failure, where uniform sequence-level rewards inadvertently reinforce hallucinated reasoning alongside correct sub-answers in complex structured outputs. Concurrently, we discover profound structural dependencies among SQL errors, revealing inherent co-occurrence semantics dictating that certain errors frequently appear together. To address the limitation and leverage this insight, we propose DiagSQL, a novel framework that utilizes an improved GRPO algorithm tailored for diagnostic Text-to-SQL validation. DiagSQL introduces Fine-grained Token-level Reward Allocation (FTRA), which parses structured responses to precisely distribute rewards at the token level, effectively isolating and penalizing erroneous reasoning. Furthermore, it incorporates Co-occurrence-aware Reward Shaping (CORS), which leverages our discovered pre-computed error co-occurrence matrix to dynamically adjust optimization objectives, encouraging logical error combinations while penalizing structural violations. Extensive experiments demonstrate that DiagSQL significantly enhances the robustness and accuracy of SQL error diagnosis, providing actionable, high-quality feedback that effectively transforms Text-to-SQL into a closed-loop paradigm. Our code is publicly available at https://anonymous.4open.science/r/DiagSQL.