PonderPounce: A Pretrained MLLM as an Episode Context Engine for Robot Control
Abstract
Multimodal large language models (MLLMs) can integrate long visual histories and infer behavior from a few examples, yet vision-language-action models rarely use this capacity as episode memory. Memory-dependent policies use purpose-built history mechanisms, while PonderPounce reuses an MLLM's native causal context as robot memory. Ponder, a pretrained System 2 MLLM, accumulates episode and demonstration context and produces continuous cognition. Pounce, a System 1 action model, asynchronously conditions control on the newest cognition and its age. They are jointly trained end to end without bridge pretraining. Optimized per-call latency meets each model's 1Hz budget, while predicted action chunks are played back at 20Hz. At RoboMME's base data scale, PonderPounce reaches 60.83% with 9B and 50.04% with 0.8B under the same Pounce architecture and interface, versus 44.51% for FrameSamp+Modul and 17.93% for \pi_{0.5} without history. With 9x data, it reaches 75.54% versus 57.88% for FrameSamp+Modul. On RoboCasa-DC, PonderPounce uses the same interface and is trained with action supervision alone, reaching 12.5% versus 11.6% for the best published demonstration-conditioned baseline and dropping to 8.6% with learned null cognition.