Dafny-to-Verus Transfer with Qwen3-14B: A Reinforcement Learning Case Study
Abstract
We study whether reinforcement learning with one program verifier improves verified code generation for another. We train Qwen3-14B with GRPO on Dafny and evaluate single-attempt verification success on held-out problem families expressed in both Dafny and Verus. Our evaluation checks for contract bypasses and we exclude unusable training specifications. On matched problems, Dafny success improves by 13.96 percentage points, compared with an estimated Verus gain of 1.23 points (95% family-bootstrap CI [0.06, 2.37]). Evidence that this Verus effect exceeds zero is sensitive to the statistical method and individual families. An exploratory policy trained directly on Verus under uniform sampling also shows only a small gain with no detectable advantage over the Dafny-trained policy. Oversampling Verus tasks with mixed outcomes improves held-out performance over uniform Verus training by 1.27 points (95% family-bootstrap CI [0.51, 2.44]). In matched 50-update continuations of the Dafny-trained policy, partial postcondition rewards provide no detectable advantage over Boolean rewards in either language. For this model and training budget, large improvements in Dafny were accompanied by limited cross-verifier gains, while targeted sampling offered a modest benefit for direct Verus training.