Frozen Memory Is Not Enough: Rethinking External Memory as Extraction
Mingyuan Li ⋅ Guangsheng Yu ⋅ Xu Wang ⋅ Shaoxiong Ji
Abstract
Methods for improving knowledge use in large language models typically fall into two regimes. Non-parametric retrieval offers flexible access to external knowledge, but adds retrieval latency, context overhead, and only shallow integration with the backbone. Parametric adaptation is efficient at inference time, but entangles knowledge with model weights and can be hard to update, audit, or transfer. Engram-style hashed memory occupies a middle regime: it stores learned information in an external, addressable table, yet consumes that table through a small learned reader. This raises a basic question: when such a memory is moved across backbones, what matters more, the frozen memory itself or the target-side reader? We study this question through cross-model frozen-memory extraction, where a memory trained on a source model is frozen and attached to a different target model while only a lightweight reader is trained. Across architectures, tokenizers, and model scales, frozen memory improves every source--target pair in a $3 \times 3$ transfer matrix, with up to 15.7\% relative perplexity reduction. Ablations show that learned memory content and correct addressing both matter, but stronger readers account for much of the transfer gain; on OpenQA, a dual-layer 4-branch reader nearly closes the gap between same-model and cross-model transfer and reaches 38.78 average. These results position Engram as a hybrid external memory and suggest that, in this regime, progress depends as much on reader design as on storage itself. Code and reproducible setups are available at https://anonymous.4open.science/r/Engram_Extractor/README.md. .
Chat is not available.
Successful Page Load