COMPASS: Statistically Verified Refinement of Probabilistic Programs with Literature-Grounded Agents
Abstract
Agentic systems that plan and execute tasks with external tools have become the de facto paradigm in software development, and they are beginning to reshape scientific workflows. Probabilistic programs are now within their reach: Language-model agents, can now revise probabilistic programs, and propose revisions. But whether a proposed revision should be accepted is a statistical question, and it should be answered by statistical evidence rather than by the agent that generated the proposal. We present COMPASS, a verifier-first system for literature-grounded refinement of probabilistic programs that automates the Bayesian workflow while separating proposal from acceptance. LLM agents diagnose models, retrieve relevant literature, and synthesize candidate revisions while a deterministic statistical verifier evaluates candidates using MCMC diagnostics, predictive comparison, posterior predictive checks, and reliability tests. Under explicit assumptions, we derive false-acceptance guarantees for predictive refinements and show that a code-only oracle cannot simultaneously achieve nontrivial power and bounded false acceptance when improvement depends on the data distribution. We demonstrate COMPASS on multiple benchmarks, and a current research problem.