On the Effectiveness of Harmsinks for Parameter-Separated Pretraining
Abstract
Removing learned information after training poses a challenging problem for AI privacy and safety, even when the culprit training data is identified. Recently emerged parameter-separated training methods do not remove undesirable data, but restricts which parameters it updates during pretraining. Promisingly these models are reported to be more resistant than typical posthoc-only unlearning to relearning certain knowledge, and may isolate sensitive copyright content and concepts. We call the designated subnetwork intended to localize information for later removal a \emph{harmsink}, which is ablated at inference time to not influence output, thereby forgetting'' the undesirable behavior. Prior work evaluates isolation using the loss gap between forget (harmful) and retain (benign) datasets; however, this behavioral metric does not establish mechanistic separation. To test whether harmful information is trulyisolated'' in the harmsink, we expand Selective Gradient Masking (SGTM) \cite{shilov2025beyond} to FineWeb data annotated for harmfulness. Across four model scales, layer-wise linear probes show that the harmful--benign distinction remains decodable after harmsink ablation, indicative that mechanistic separation often failed to emerge despite loss gaps. Our findings suggest that forget--retain loss differences can overstate parameter separation and motivate further work toward more reliable harmsinks.