Stress-Testing Geometric Interpretability Claims in Grokking
Abstract
In this paper we test whether Hodge decomposition gives a consistent account of how a transformer’s learned representations change as it groks modular addition. The generalising solution to modular addition has an independently understood Fourier structure, giving us a known representation against which to compare our geometric measurements. We measure how the same probe activations move between checkpoints and separate this motion into exact (gradient-like), coexact (curl-like), and harmonic components. The baseline run suggests coexact-led motion during the transition followed by exact-led motion after grokking, but this pattern does not appear in every run. The consistent result is a relative shift towards exact motion. From the transition to the post-grokking phase, the difference between the coexact and exact fractions decreases across all 56 combinations of eight training seeds and seven analysis settings, with uncertainty calculated across the eight independent seeds. At the baseline setting, the mean change is −0.267, with a 95% t interval of [−0.400, −0.135]. The result remains consistent when we resample the probes and assign checkpoints to phases separately for each seed. When we permute the probe identities between checkpoints, the higher exact fraction observed after grokking disappears. Lastly, the method recovers the expected components on fields with known Hodge decompositions.