EnergyGradMem: Learning an Energy Function for Test-Time Memory Writing
Matvey Kairov ⋅ Ivan Rodkin ⋅ Mikhail Burtsev ⋅ Yury Kuratov
Abstract
Compressing context into a compact memory state lets a language model answer future queries without retaining or reprocessing it. GradMem writes such memories by test-time optimization of a context-reconstruction objective chosen a priori. We introduce EnergyGradMem, which instead learns a query-independent loss-energy function from downstream supervision. At inference, it writes to memory by minimizing the loss-energy over a context, and uses only memory and query for inference. In multi-query associative recall (MQAR) experiments we compare learned loss-energy with reconstruction-based memory update, refinement and longer-context generalization.
Chat is not available.
Successful Page Load