From Vibe-Coding to Verified-Coding: A Case Study in Human-Architected, AI-Assisted, Human-Audited Verification of MoE-Inference Support in nano-vLLM
Abstract
AI4Code models have made “vibe-coding” routine, yet the correctness of AI-generated code still rests on testing, not verification. As agents grow able to draft proofs cheaply, a division of labor emerges: a domain expert architects the system and states each subcomponent’s properties, an AI agent drafts the proofs, a human expert audits them against a composition theorem, and an agent assembles the verified components end-to-end. We instantiate this workflow in a brownfield case study—adding Qwen3-30B-A3B Mixture-of-Experts (MoE) inference to nano-vLLM, a pedagogical vLLM fork—which requires composing grouped-query attention (GQA), tensor parallelism (TP), expert parallelism (EP), data parallelism (DP), a fused-MoE Triton kernel, and machine-checked contracts for their composition. Our contributions are threefold. (1) A verified microbenchmark suite: nine hand-implemented subcomponents, each with an AI-drafted Verus contract covering shape, dtype, and routing/sharding invariants, which Verus machine-checks and a human audits against the implementation; released open-source. (2) A composition theorem: the composed MoE variants refine a common reference semantics in the Verus model and satisfy four global properties—completeness, disjointness, data-race freedom, and deadlock freedom. It certifies two functional-block swaps: a fused Triton kernel for the per-expert FFN loop, and a single all-reduce for three all-to-alls in the EP schedule at dp=1. (3) An agent-composed integration: given the components and their contracts, a coding agent builds the end-to-end nano-vLLM model, matching vLLM’s GSM8K accuracy. Swapping in the certified-equivalent variants yields up to 1.58× end-to-end speedup—1.39× from the fused Triton kernel and 1.14× from the single all-reduce. Together, these results provide a template for using AI models to build verifiable critical-system software, such as model inference infrastructure.