Better Language Models Require Better Domain-Specific Inductive Biases
Abstract
Language models (LMs) obtain much of their capability from the diversity of their training data. This is possible because neural architectures like transformers are remarkably flexible and can handle a variety of domains such as natural language, code, and mathematics. This raises the question: are transformers optimal for any specific domain? In this position paper, we present empirical evidence that the inductive biases of transformer-based LLMs should be re-examined, since, while broadly useful, they are often suboptimal on subsets of their training data. Methods. To support this position, we use a method that learns new activation functions within a transformer as a tool to explore the space of inductive biases. We train the resulting modified architectures on various data subsets. A large-scale evaluation shows that vanilla transformers are relatively easy to improve, but mostly on a per-domain basis. For example, an architecture optimized for a natural language dataset like FineWeb performs well on another such dataset like TinyStories, whereas architectures optimized for mathematics perform poorly on natural language. We also identify the mechanisms responsible for these improvements. Implications. These results call for re-evaluating core design choices in language models. We do not advocate a return to handcrafted, domain-specific models. Instead, we emphasize that different domains can benefit from different learning mechanisms. This will require new designs that preserve cross-domain interactions, analogous to how reasoning in the brain arises from the interaction of multiple specialized networks. We argue that this perspective may be key to improving reliability and data efficiency.