Can Data Attribution Filter Out Subliminal Learning? Not Reliably.
Abstract
Subliminal learning allows language models to transmit behavioral traits through training data with no obvious semantic relationship to those traits, undermining content-based data filtering as a safety intervention. Training data attribution offers an alternative: it identifies influential training examples from model gradients rather than from content, and so may apply in exactly the cases where semantic inspection fails. We evaluate three attribution methods for this purpose (GradCos, a contrastive GradCos variant, and EK-FAC) against divergence tokens, a strong oracle baseline previously shown to localize subliminal learning, across three models and multiple target preferences. Filtering at the token level, EK-FAC mitigates a significant part of the effect, the other methods provide little benefit, and all fall short of divergence tokens. Filtering entire samples is less effective for every method, though EK-FAC often gives a stronger signal than divergence tokens in this setting. Success is inconsistent across methods and settings: variants that work well for some model–preference combinations fail for others, and we do not identify a consistent explanation for these differences. Our results suggest that gradient-based attribution can identify data responsible for subliminal learning in some settings, but that some approximations are more reliable than others.