Aligning Large Language Model Agents with Rational and Moral Preferences: A Supervised Fine-Tuning Approach
Amit Dhanda ⋅ Daniel Chen ⋅ Christian Hansen
Abstract
LLM judges are used where no ground truth exists, which is also where their reliability cannot be checked. We argue that progress on judge reliability needs the opposite setting---tasks with a cheap, model-independent optimum---and that canonical economic games supply one: the optimal policy under an explicitly stated utility function is computable by a solver, so a judge's verdict can be scored against a target no model produced. We use this setting to make three points about automated evaluation. First, a solver-grounded judge surfaces failures that outcome-level scoring misses: off-the-shelf LLM agents not only deviate from payoff-maximizing play but hold beliefs inconsistent with their own actions, an incoherence invisible to any metric that sees only final answers. Second, judge design changes conclusions: scoring every component of a response conflates genuine error with arbitrary tie-breaking, and we show that restricting credit to components whose value actually changes utility separates a learned decision rule from a memorized template---the same model moves from $0.55$ to $0.85$ depending only on which criterion is applied. Third, the grader doubles as a supervisor, and an ablation shows that training on its worked derivations rather than its final labels is what transfers the computation. We offer this as a testbed where claims about judge reliability can be settled by construction rather than by appeal to a stronger judge.
Chat is not available.
Successful Page Load