CondenseVLA: Learnable History Condensation for Efficient Multi-Frame VLA
Abstract
Robot manipulation often depends on short-term cues—such as brief contacts, temporary occlusions, and intermediate object states—that are ambiguous in a single frame and can only be resolved by integrating observations over time. However, injecting historical information solely by enlarging the history window is computationally costly and frequently yields limited gains due to substantial redundancy and task-agnostic, inflexible sampling of observations. We propose CondenseVLA, a multi-frame VLA policy with a lightweight learnable history condensation module for using short-term visual context in action decoding. Given a sequence of cached per-frame features from the VLM backbone, CondenseVLA distills observation history into a fixed set of slot tokens using learnable queries and multi-head cross-attention. Simple auxiliary regularization encourages diverse attention patterns by penalizing overlap among slot attention maps and reduces slot 12 redundancy by decorrelating slot embeddings, stabilizing learned condensation without imposing strong time priors. By expanding the history window and distilling it into a compact token set, CondenseVLA improves over the baseline on the evaluated Simpler-WidowX tasks, increasing success rate from 58.4 to 70.8 (+12.4), achieves 97.2% on LIBERO, and reaches 52.0% on real-world tasks. Beyond the gains, our results show that long history windows contain substantial redundancy, so naively providing more frames yields limited benefits, whereas learned condensation into fixed slot tokens provides an effective bounded-token way to leverage temporal context.