Remembering What Matters: From Markovian to Subtask-Causal Memory in VLA Policies
Abstract
Long-horizon robotic manipulation is often non-Markovian: the correct action can depend on earlier state changes that are no longer visible in the current observation. Vision-language-action (VLA) models provide strong generalist robot policies, but their performance can drop when a task requires remembering which earlier subtasks changed the world. Existing memory mechanisms extend context through recent-frame windows, compressed latent stores, or similarity-based keyframe retrieval. They do not explicitly distinguish which completed subtasks are relevant from which frames within those subtasks carry the needed evidence. We introduce TTC-MEMORY, a two-tier memory framework for long-horizon VLA policies. The first tier segments execution into subtasks and constructs a dependency directed acyclic graph (DAG), so retrieval is restricted to ancestors of the current subtask. The second tier adaptively stores state-change keyframes within each retained subtask and injects them into a frozen VLA backbone through graph-aware cross-attention, with fallback to recent history when planning or clustering is uncertain. We evaluate TTC-MEMORY on LIBERO, RMBench, SimplerEnv-Bridge, the SafeLab chemistry-lab simulation benchmark, and Franka real-robot tasks. The comparisons use a shared VLA backbone, training budget, simulation seed protocol, and physical-trial protocol. TTC-MEMORY improves the long-horizon and memory-dependent slices in this matched setting. On the 10-subtask SafeLab slice, it reaches 58.3% success, compared with 38.2% for the strongest matched prior memory baseline, MemER, and 45.1% for DAG-only retrieval. Real-robot trials provide supporting transfer evidence rather than the main claim. Ablations show that the two tiers are complementary: DAG retrieval increases long-range dependency recall, while adaptive keyframe selection removes redundant history and preserves attention on task-critical state changes.