Neural networks are more modular than single neurons suggest
Abstract
Neural networks routinely develop representations in which individual neurons respond to mixtures of task-relevant variables, leading to the widespread view that such networks lack modular organization. We challenge this conclusion by introducing a formally nested hierarchy of modularity: unit-level, which requires distinct computations to occupy disjoint neurons; orthogonal, which asks whether a rotation reveals functionally independent subspaces; and linear, which asks whether a bounded linear map can separate task-relevant computations into independent subspaces. We develop a unified framework for learning decompositions at each level and evaluating them causally through selective ablation. Across vision models, anguage models and multitask RNNs, we find that networks which lack unit-level modularity nonetheless often exhibit clear functional modularity at the subspace level. In vision and language models, content versus style and syntax versus semantics respectively dissociate under subspace ablation despite sharing overlapping populations of neurons. In RNNs trained on multiple cognitive tasks and previously shown to lack modular structure, our decomposition recovers sharp double dissociations between tasks. These results demonstrate that modularity cannot be assessed at a single level of description. The hierarchy we introduce provides a principled framework for determining when and at what level of analysis functional specialization exists in neural network representations.