GRID: Grammar-Railed Decoding for Verifiable, Auditable Constrained Generation
Abstract
Deploying LLM code generation where correctness is contractual—enterprise SQL under role- and schema-level policy is our driving case—needs outputs that are valid and policy-compliant by construction, guarantees whose assumptions are explicit, and a record an auditor can check. We present GRID (Grammar-Railed Decoding), a neuro-symbolic constrained-decoding engine that keys exact next-token masks on parser configurations (lexer scan state × LALR(1) stack) rather than on token sequences, and uses the incrementally advanced parser as a viable-prefix oracle. A byte-level trie walk bridges LLM tokens to grammar terminals, and a context-independent/context-dependent split makes cache-key soundness hold by construction. Role-based access control is compiled into the language: role projections subset the grammar's productions and schema lexicons restrict identifiers, so forbidden verbs and tables are unreachable at mask level. Soundness, completeness, termination, and flat per-token cost are stated with explicit preconditions, and we separate what is proved, what is assumed and checked, and what is only measured. On the full Spider dev set (n=1,034, paired 95% intervals) the mask raises execution accuracy by 13.7 points at 0.5B ([10.7, 16.6]), and by 11.6 on the held-out test databases. At 7B its effect is within noise on dev and slightly negative on test, and the 2.5-point gain of GRID with one checker-triggered repair round is matched by the same round on unconstrained output: for capable models the mask buys policy, not accuracy. With one forbidden table per database, a prompt-level instruction leaks that table in 86% of the 7B answers that need it and hiding it from the prompt in 22%; the mask leaks in none of 2,068 outputs, at no accuracy cost on the other questions. Column-level policy is outside exact mask reach: with unbounded alias names the compliant language is not context-free, so GRID checks it post-parse and retries. Rust kernels give a 3.6–6.7 µs median per-token mask (llguidance keeps the better p99), and a hash-chained per-token audit trail replays bit-identically. GRID verifies syntactic and policy correctness, not functional correctness. Code: https://github.com/evolutionIdGmbH/grid.