What Is Lost in Post-Training Language Models? Maintaining Access to Pluralistic Perspectives
Abstract
Previous work has demonstrated that classic post-training methods can cause minority perspectives to disappear from the trained policy, but has focused on the model's behavior under neutral prompting. We argue that this collapse is not restricted to the trained policies' default behavior, but occurs in several distinct modes, which are important to distinguish. Default collapse: the policy strongly concentrates on one perspective under neutral prompts. Recognition loss: the model can no longer recognize whether a presented response is consistent with a specified perspective. Steerability loss: the model can no longer act according to a perspective even when explicitly instructed to. These three modes can occur independently and have different practical implications: a concentrated default may be acceptable, whereas steerability loss removes the possibility of representing minority perspectives, even when they are explicitly requested, rendering the policy clearly non-pluralistic. We demonstrate empirically that all three modes arise under standard post-training, and that recognition loss and steerability loss are closely coupled. Building on this, we propose recognition tuning, a supplementary training phase in which the model is shown recognition prompts, and show that it recovers both recognition and steerability while leaving the model's trained default behavior largely unchanged. To counter default collapse, we present a new method for distributional alignment, which we term stance-distribution matching: we extend the KL-regularized post-training objective with a second Kullback-Leibler term that penalizes the divergence between the distribution of perspectives the policy expresses, as scored by an LLM judge, and a prespecified target distribution. By operating on perspectives rather than responses, it is robust to semantically equivalent reformulations of responses. We show that the method is able to shift a given policy toward a pluralistic target distribution and that this behavior generalizes to held-out scenarios.