When Correctness Is Computable: Auditing AI Evaluators Against an Exact Oracle in Household Finance
Vignesh Coumarane
Abstract
Two problems make AI evaluation hard to trust: benchmarks leak into training, and we usually have no ground truth, so we validate a learned evaluator against noisy human labels. Both problems disappear in a domain where correctness can be computed. We study one: allocating a household lump sum, where a plan is valid only if it obeys three hard invariant - the parts sum to the whole (conservation), a free employer match is not left on the table (dominance), and cash is not parked idle when a return exists (efficiency). A twenty-line oracle decides validity exactly, over scenarios generated on the fly, so there is nothing to memorize and no annotator to disagree with. We use this oracle to audit the two evaluators practitioners reach for first, an LLM judge and self-consistency, and to ask a question the usual setting cannot: not whether they fail, but by how much, on which rule, and whether their own confidence would warn a deployer. A cheap judge panel accepts 56.7% of invalid plans when asked plainly. Stating the invariants in the prompt removes almost all of it; making the judge compute the checks first adds nothing measurable ($p = 0.45$). Majority voting does not help, because the errors are systematic rather than random. And the judge's confidence is no guide to its correctness - the plans it accepts by mistake carry confidence equal to or higher than the ones it rejects correctly. Where a property is invariant-checkable, a free deterministic oracle beats the learned evaluators on accuracy, cost, and auditability at once. We release the testbed.
Chat is not available.
Successful Page Load