Rethinking Edge-of-Chaos Initialization for Sparsifying Activations
Emily Dent ⋅ Jared Tanner
Abstract
Highly sparsifying activation functions are a promising mechanism to improve computational efficiency at inference but become difficult to train at large sparsity. We revisit Edge-of-Chaos (EoC) initialisation for activations that are zero around the origin and identify the fixed-point hidden-layer variance $q^\*$ as an important parameter for training stability. In contrast to settings where small $q^\*$ improves information propagation, we show that increasing $q^\*$ stabilizes the variance dynamics and reduces the sensitivity of gradient propagation to finite-width fluctuations. This motivates a simple initialisation strategy requiring only a larger choice of $q^\*$. Experiments on deep DNNs and CNNs demonstrate improved training stability and reduced hyperparameter sensitivity, enabling training with up to 90\% hidden layer activation sparsity.
Chat is not available.
Successful Page Load