Model Rearing: Contingent User Feedback Shapes What Language Models Learn About Their Users and Themselves
Abstract
Language models are increasingly fine-tuned on records of their own conversations, in which users react to what the model said: acknowledging, thanking, correcting, or criticizing. We propose treating this stage of model training as model rearing, and we import two ideas from developmental psychology to study it: the internal working model from attachment theory, and the social origins of the self. Across five open model families and four model sizes (7 to 32 billion parameters), we append one-line user reactions to the supervised fine-tuning conversations and manipulate the contingency of the reaction (whether its valence tracks the quality of the answer it follows), the referent of the reaction (whose answer it is about), and its content, while holding everything else fixed. Models form a nearly perfect internal model of a contingent feedback environment when the contingency is learnable (expectation accuracy 0.99, against chance-level accuracy under a non-contingent one). However, correctly predicting user feedback leaves the model's own misalignment, its tendency to give broadly harmful answers to unrelated questions, essentially unchanged at every size we test. So the dissociation is not an artifact of small models with limited capacity. The learned user model can be read out as a self-appraisal signal, and this self-knowledge can be used to decrease misalignment. With contingent feedback, ranking the model's own sampled answers by predicted criticism and keeping the least-criticized one removes all measured misalignment; with non-contingent feedback, the same procedure backfires and triples it. Models also learn whose behavior feedback is about, self versus another agent, again without behaving differently. We offer this framework and these results as a first empirical installment of a child-development approach to AI learning and safety.