Memory Grafting: Scaling Language Model Pre-training via Offline Conditional Memory
Abstract
Large language models acquire rich internal knowledge during pre-training, yet compact models are still typically forced to relearn such knowledge from raw data or imitate it indirectly through distillation. We investigate a different route: can pretrained knowledge itself be directly reused during pre-training? To this end, we propose Memory Grafting, a representation-level transfer paradigm that extracts structured memories from a large frozen donor model and injects them into a compact student throughout training. Unlike knowledge distillation, which transfers behavior through output alignment, or retrieval-based methods, which expose external knowledge only at inference time, Memory Grafting makes pretrained knowledge part of the student’s optimization process itself. This suggests that knowledge in large language models can be modular, transplantable, and reusable across scales, pointing to a new path toward more efficient language model pre-training.