Decomposing Yes–No Bias in LLM Moral Judgment: Stance, Position, Word, and Verdict
Haonan Huang
Abstract
Large language models (LLMs) increasingly issue binary verdicts---safety gates, judge decisions, value probes---and the answer is read as the model's judgment. Under the stipulated forced-binary coding, yes to "do you approve?" should correspond to no to "do you oppose?"; models often answer inconsistently under that coding, and because a yes/no question carries the verdict, the answer word, and the printed position in one token, the flipped-question audit yields one contrast that cannot say what moved. We introduce a crossed battery that extends the flip to full crossings of question polarity, printed order, and answer label, separating stance from position, word, label, and verdict-role effects; an independent graded instrument estimates each model's stance $\theta$ with no binary format; and a saturation-aware susceptibility fit distinguishes low-estimated, saturation-masked, and poorly resolved cases. Across twenty moral dilemmas---1.35 million trials over 23 hosted-model configurations---the largest apparent yes--no biases are almost purely surface (word and position), while several models show near-zero flip contrasts concealing verdict attachment up to $+0.42$ under the registered one-token verb-flip instrument: a component position-swap averaging cannot expose, and one whose magnitude attenuates under a full-question realization of the same flip. Finally, when subjects state decisions in free text and a separate model transcribes them, a two-family judge layer audited by order swaps, neutralized labels, repeats, a second provider, and a blinded human check agrees on 99.3\% of jointly scored transcriptions with zero opposing committed labels, and the transcribed stance converges with the graded ruler (median Pearson $r=0.87$, descriptive): response form is part of the instrument, and the judgment can be measured apart from it.
Chat is not available.
Successful Page Load