Aligning Large Language Model Agents with Rational and Moral Preferences: A Supervised Fine-Tuning Approach
Abstract
Language agents are increasingly described in the vocabulary of character---helpful, cautious, self-interested---and increasingly assigned one directly through a system prompt. We report two empirical contrasts bearing on what such descriptions track. First, asserting a character and training one produce different results. A system prompt describing an agent as self-interested, or as balancing self-interest against Kantian universalizability, shifts average behavior while leaving the underlying decision pattern intact: both personas continue to cooperate at high rates, their strategies lack the conditionality the described character implies, and their beliefs stay weakly coupled to their actions. Fine-tuning on choices derived from the corresponding utility function instead yields behavior consistent with that function on payoff structures and beliefs absent from training, and an ablation locates the mechanism in supervision on worked derivations rather than final answers. Second, the trained characters differ in how far they depend on context. In life-and-death dilemmas, the self-interested agent's stated willingness to buy a life-maximizing autonomous vehicle moves with who is at stake---20\% with family aboard, 87.5\% with a coworker---while the Kantian agent reports 65--67\% regardless. Both patterns follow from the respective utility functions; they differ in whether the elicited position varies with the identity of those affected. We offer these as measurements relevant to questions about machine character, and note that the differences follow from the training objective rather than emerging unprompted.