Where Memory Belongs: LEDGER, an Object Ledger for Memory-Augmented VLAs
Tanguy Dieudonné ⋅ Jack B Jedlicki ⋅ Heng Yang
Abstract
Memory is essential for long-horizon, partially observed robotic manipulation: a robot must remember which object was placed in a drawer, whose cup it moved, or how many action cycles have elapsed. Recent vision-language-action (VLA) models embed memory directly inside the policy, but benchmarks show no single in-policy mechanism covers all spatio-temporal dimensions, trailing oracle methods by a wide margin. We argue that memory type dictates where memory should reside: short-term perceptual memory (repetition, timing, retracing) belongs inside the policy, while long-term object memory (persistent spatial state, containment, event history) belongs outside as an explicit, readable record. We present Ledger, a harness that realizes this split over a single fine-tuned $\pi_{0.5}$ policy by pairing an in-policy frame-sampling memory with an external spatio-temporal object memory, the ledger (a SAM3 tracker, a VLM captioner of the demonstration, and an LLM that decides at step boundaries). On RoboMME, Ledger reaches the best performance reported on the benchmark, a $64.0\%$ four-suite average (vs. $45.9\%$ for the strongest prior method under identical evaluation), leading object reference ($58.0\%$ vs. $40.3\%$) and object permanence ($86.8\%$ vs. $56.2\%$) using a single set of weights. Choosing the memory source at runtime, from the instruction and the record, removes the need for a task-level router.
Chat is not available.
Successful Page Load