Justified Autonomy: Decision-Level Warrants for Meta-Agent Harnesses over Authority-Signed Knowledge Property Graphs
Abstract
Meta-agents design, train, and supervise other agents. To supervise a worker, a meta-agent needs to know whether a particular answer is right, but the guarantees we have for learned systems do not work that way. Benchmark accuracy, calibration, and conformal coverage are all statements about a population of decisions, and none of them says anything about the decision in front of you. In this paper we give the supervisor a contract on each decision instead. The world model is a knowledge property graph, built from typed claims that a model reads once from an authority's documents; each claim records who stands behind it and whether it still stands, and we certify the claims against the source bytes before any decision uses them. Decisions are made by honing, which eliminates candidates only where the authority's own stated precedence licenses it, and each commitment ships a witness that replays without the model that produced it. What the graph does not settle becomes an abstention that names its cause. We measure two arrangements. With no model at decision time, the system answers all 262 probes of a five-hazard register correctly after one certified remediation, while reading a tenth of the corpus per decision; under 90 injected intake faults the engine alone commits 1,236 silently wrong decisions, and a certification ladder reduces that to zero. With a model deciding and the graph checking it, which is what a deployed harness runs, a small model is wrong five times on held-out cases of a federal regulation; after we revoke every decision the graph does not license, it is wrong on none of them, at 38\% coverage. We also measure where the approach fails: on a register whose controlling authority states no rule for its central question, a graph built only from what the authority states abstains on 94 of 95 cases, and the only graph that decided anything was one whose rule we had written ourselves.