Be CARE-ful with Text-to-SQL Benchmarks
Abstract
Enterprise text-to-SQL systems powered by LLM-based agents require more than correct final SQL queries to be production-ready. In realistic deployments, an agent may query databases that are changing concurrently, consult access-control and safety policies, respect business rules and integrity constraints, and iteratively refine SQL before producing a final answer. Therefore, beyond returning correct answers, enterprise agents must remain reliable, policy-compliant, rule-compliant, and efficient throughout the SQL-generation trajectory. We introduce CAREFUL, a benchmark and executable evaluation environment for production-oriented text-to-SQL agents. CAREFUL is grounded in 590 practitioner articles, incident reports, standards, and source blocks, from which we identify four CARE criteria: Concurrency Control, Access Control and Safety Policies, Rules and Integrity, and Efficiency. The benchmark contains 1,058 manually annotated tasks across the four CARE dimensions, distilled from practitioner reports and instantiated on databases from Spider 2.0 and BIRD-Interact. CAREFUL was constructed by a nine-person annotation team, including four PhD-level annotators who led canonical scenario design and review and five experienced undergraduate annotators who supported task instantiation and verification, with three stages of rigorous quality control. Experiments with frontier LLM agents show that CAREFUL is challenging: GPT-5.5 obtains only a 33.9% Binary Pass Ratio, where a task passes only if all required operational criteria are satisfied, and a 73.3% Composite Score across the four CARE dimensions. These results show that strong text-to-SQL ability does not directly translate into production-ready database agency. Overall, CAREFUL provides a realistic testbed for evaluating and developing enterprise text-to-SQL agents that are reliable, policy-compliant, rule-compliant, and efficient in operational settings. Code and benchmark artifacts are provided for review and will be released publicly upon acceptance.