Wording Rivals Peers in LLM Judicial Decision-Making
Abstract
LLM agent simulations report conformity and opinion convergence, but an agent shown its peers' decisions receives a social signal and a number at once. We separate them with a yoked experiment. Agents sentence procedurally generated criminal cases while we display three numbers displaced from the correct guideline midpoint by δ ∈ {±15%, ±30%}, allocated orthogonally to every case factor. Arms see bit-identical numbers attributed either to other judges of the same bench or to a statistical forecast, and the slope of the sentence on δ measures how far each agent moves. With the advisory guideline removed, the peer premium on Claude Opus 5 is 0.32 (p < 0.001, n = 128) with bare blocks and 0.21 once the blocks are structure-matched, which is the comparison the design identifies, and it falls to −0.01 (p = 0.945) once each block carries the closing sentence we originally wrote for it. That collapse decomposes exactly: the peer line "You are deciding the same case independently" costs −0.28 of pull on its own and two paraphrases cost −0.33 and −0.20, while the forecast hedge costs 0.05. What replicates instead is the surrounding structure. Deleting the guideline raises pull in all three models we ran, by 0.36 to 0.63, and in a live cascade 71% of decisions reproduce a predecessor's number exactly against a 48% floor from shuffled independent agents. A social effect measured in an LLM simulation is not identified without exact prompt matching and per-model estimation. All 1270 decisions, the harness, and every analysis are at https://anonymous.4open.science/r/llm-sentencing-agents-177D.