Remembering Where Things Are: Measuring What Persistent Spatial Memory Is Worth — and What It Costs — on a Physical Robot
Chi Feng Chang
Abstract
Embodied agents built on foundation models typically perceive each task from scratch: what the robot saw a minute ago is gone. We present MEM, an open-source persistent spatial-memory layer that ingests RGB-D and pose from a commodity phone, maintains a queryable 3D object memory, and serves it to agents over REST/MCP. We measure the task-level value of persistence with a pre-registered ablation on a physical wheeled robot: an LLM-driven agent navigates to language-specified objects either (a) by querying memory built from a one-time scan, or (b) memoryless — restricted to observations made during its current in-place observation round, so nothing persists across standpoints and the agent must physically search. The two conditions share perception, pose, controller, and an image-verified target-selection rule; only access to the past differs. Across 60 pre-registered trials in five layouts, persistence did not buy completion — both arms arrived in 26/30 trials, with tape-measured completion 83% (memory) versus 87% (memoryless) at the pre-registered 50 cm threshold and failure counts symmetric at four per arm — it bought time: median time-to-arrival 34.6 s versus 65.8 s (one-sided Mann–Whitney $p<10^{-4}$ with all failures assigned a common 900 s horizon; memory won the pair majority in all five layouts, exact sign test $p=0.031$), with a per-object gradient from 1.4× on easy objects to 8× on the hardest — persistence is worth most exactly where re-detection is hardest. The price is a one-time handheld scan of at most three minutes per environment — charged in full against memory in our accounting, and amortized after roughly two to six tasks. Code, per-tick trial logs, and the operating protocol are released.
Chat is not available.
Successful Page Load