Synth-Ethics: Moral Outlook as a Controlled Variable in Synthetic Training Data
Hanna Shcharbakova ⋅ Pieter Delobelle ⋅ Anastasia Stasenko ⋅ Carlos Rosas-Hinostroza ⋅ Pierre-Carl Langlais
Abstract
When a language model stands in for a person, as a simulated user, an annotator, or a judge, it enacts a moral outlook that nobody chose: whatever its training distribution happened to contain, or whatever a prompt can induce. What that stance is can only be guessed at: even when the intended outlook is written down, as for instance with Claude's constitution, what the trained model actually enacts on a given dilemma stays unknown. We propose Synth-Ethics, where we make that outlook an explicit, controlled variable. We operationalise $16$ sub-frameworks from six ethical traditions as framework cards, generate 68,520 framework-conditioned reasoning traces over 4,397 dilemma scenarios built along two routes, reshape them into the format of an open generalist synthetic corpus, and train one small model per framework, so that the outlook is carried by the weights rather than by a prompt. The result is a population of moral-outlook simulators: small open models that each enact one documented evaluative stance and share everything else. To test the effect of our framework cards during post-training, we create four single-framework finetuned models of Qwen3 $0.6$B answer 250 held-out dilemmas. We find that each model's own framework ranks first or second among the sixteen candidate reference policies. Furthermore, the same descriptions given in the prompt recover a fraction of that agreement on an instruction-tuned control, drift toward one consequentialist neighbourhood whichever framework is named, and elicit no classifiable verdict from the base on $93\%$ of items. Synth-Ethics offers simulation research a way to state explicitly whose morality a synthetic user enacts.
Chat is not available.
Successful Page Load