SQL-Zero: Self-Evolving Text-to-SQL
Abstract
Training a competitive Text-to-SQL agent usually depends on human-annotated natural-language/SQL pairs, which are expensive, domain-specific, and a bottleneck for scaling to new databases. We show it is possible to train a competitive solver with zero annotated pairs. We introduce SQL-Zero, a proposer--solver self-play in which a challenger and a solver start from the same base LLM and the only ground truth is execution against the database itself. The challenger generates SQL pairs calibrated to the solver's current difficulty (targeting "hard but solvable"), and both roles are updated with GRPO in alternating turns, with a template-level repetition penalty on the challenger to prevent diversity collapse. Training on BIRD databases with no labels, self-play improves over the zero-shot base on BIRD dev by 5.1 points at 3B and 6.3 points at 7B, and exceeds a Spider-gold control trained under the same recipe. It also scores higher than a matched control trained on human BIRD gold over the same databases, exceeding it under an exact paired test. Finally, self-play does not overfit its training domain: it still outperforms the base on Spider and, under lexical perturbation (Spider-Syn), degrades less than the matched BIRD-gold control.