The Price of an Explanation: What Demanding Verifiable Artifacts Costs, and What It Buys Back
Edgar A Duenez-Guzman ⋅ Eduardo Dueñez ⋅ Edgar Said Hernández Sánchez ⋅ Suzanne Sadedin
Abstract
Generative models produce many candidate solutions but no intrinsic signal by which to rank them. One mitigation is to demand a verifiable artifact: a program that can be run against the available evidence. The standing objection is that this incurs costs, both in solution quality and computation budget. We measure these two costs on the $120$-task ARC-AGI-2 evaluation set. Demanding a verifiable artifact from an LLM reduced the score from $93$ to $68$ of $120$ tasks ($77\\%$ to $57\\%$). At least $21$ of those $25$ tasks are the price of the explainable form itself, reached by no program-producing configuration we ran. A short evolutionary search step over two already-sampled programs, at no further model cost, beats best-of-$2$ sampled code by $9.2 \pm 1.4$. Demanding verifiable artifacts also reduced computation spend, relative to a sampler given the same budget, by $15\\%$-$49\\%$ depending on the budget, because the artifact enables adaptive stopping when a problem is solved. We tested the same composition of verified artifacts with evolutionary search in quantum error correction, showing that this methodology can expand to real-world scientific applications.
Chat is not available.
Successful Page Load