MathaMAGICal: Validating Data Attribution during Arithmetic Acquisition
Abstract
We study when trajectory-based data attribution predicts retraining effects during arithmetic acquisition. GPT-2-style models (86M and 152M parameters) train from scratch on web text and four matched arithmetic data families, then compose two operations in an order withheld from training. Using MAGIC, we differentiate held-out loss through the realized training trajectory and test first-order predictions against data-removal retraining. Full-run attribution at 86M overflows; clipping restores finite scores but leaves weak retraining agreement. Short windows around acquisition validate at both scales, including four windows on a 152M model that solves all 200 held-out problems after training on 3.0B tokens. Within validated windows, primitive facts and worked examples are enriched in both score tails, indicating sensitivity in both directions. Family-level patterns replicate across configurations despite unstable individual block rankings. At 152M, matched deletion reproduces the predicted ordering among the four arithmetic families; web deletion also changes the arithmetic fraction and is interpreted separately. These findings identify a regime for measuring local data sensitivity during capability acquisition.