O(3)-Equivariant Sparse Autoencoders for Neural Network Potentials
Mannat V Jain ⋅ Joy Z Yang ⋅ Ziyu He ⋅ David Olloqui ⋅ Ward Gauderis ⋅ Thomas Dooms
Abstract
Sparse autoencoders have proven useful for interpreting neural network activations by decomposing them into sparse, learned features. One common architecture is the TopK sparse autoencoder, which preserves the $k$ largest latent activations for every sample and zeroes the rest. Since molecules can be rotated arbitrarily in three-dimensional space, rotating a molecule should not change invariant property predictions and should predictably transform equivariant ones. Equivariant chemical models transform groups of hidden activations according to known rotation rules, but standard sparse autoencoders trained on these activations incorrectly treat rotating-group coordinates independently. A standard sparse autoencoder thus assigns different features to rotated copies of the same molecule, making feature interpretation dependent on an arbitrary coordinate frame. We introduce O(3)-equivariant sparse autoencoders, which preserve or discard entire rotating groups at a time and rank them according to rotation-invariant magnitudes. On held-out MACE-MP-0 activations, this approach achieves explained variance 0.9992 with rotational symmetry enforced by construction. Removing a learned feature selectively changes predicted barriers for its matching bond family, demonstrating that the feature participates in MACE's energy computation. This intervention distinguishes bond-change families with AUROC 0.806 compared with 0.593 for equivariant PCA.
Chat is not available.
Successful Page Load