Repeated retraining can weaken LLM distortion of collective opinions
Abstract
By selecting, filtering, and framing information, AI models can influence individual opinions. Through peer interaction, this influence can propagate across populations, shifting collective opinions towards model-specific outcomes. In this work, we investigate how repeated retraining of AI models affects this process. We conduct a large-scale simulation study deploying a dozen open weight models within an established model of human–AI co-evolution, and measure model influence as the deviation of collective opinions at equilibrium from those that would emerge in the absence of LLM influence. Fine-tuning on population responses can reduce this deviation, whereas with in-context learning this effect can vary. The strength of regularization towards the pretraining policy controls the magnitude of this effect: weaker regularization leads to equilibrium opinions closer to the baseline without LLM influence. Overall, our results suggest that by adapting models to the population--rather than populations adapting unilaterally to the model---a model's influence on equilibrium outcomes can be reduced.