Developmental Adaptive Attention Masking for Transformer Training
Abstract
Full global attention exposes Transformer layers to long-range dependencies before useful local representations have formed. We propose Developmental Attention Masking (DAM), a curriculum that begins with local attention and expands each layer’s receptive field using normalized attention entropy. Unlike fixed sparsity or hand-designed schedules, DAM adapts expansion timing layer-wise. Our analysis interprets entropy as within-window selectivity, relates it to local softmax sensitivity, and identifies data-dependent conditions motivating expansion. Experiments on structured sequence tasks, character-level language modeling, and autoregressive models up to 6.9B parameters show improved token efficiency and reduced active-attention cost with comparable or improved final performance. Motivated by developmental “starting small,” DAM provides an adaptive curriculum over attention capacity.