Alignment Needs 'Cognitive Control': On The Role of Regularization in LLM Alignment
Abstract
This position paper argues that reference-anchored regularization in large language model alignment should be understood not as engineering overhead but as a functional necessity, the computational equivalent of cognitive control in the brain. The alignment community increasingly frames the cost of regularization as an "alignment tax" to be minimized, pursuing reference-free methods that remove the pretrained reference policy. We argue that this orientation is misguided. Cognitive neuroscience has established that the cost of overriding default behavior is a functional feature of intelligent systems, preventing overcommitment to noisy reward signals and preserving prior behavioral repertoires; its impairment is a causal vulnerability factor for addiction. When formalized as a divergence penalty from a structured default, this cost yields the same regularized objective used in alignment, and the pathologies observed when it is removed, including catastrophic forgetting, reward hacking, and loss of generalization, correspond to the consequences of impaired cognitive control. Rather than minimizing the cost of regularization, the field should design its structure systematically, specifying what the reference encodes, how the penalty on deviation should vary across contexts, and how to evaluate what the regularization preserves rather than only what it costs.