Rethinking Causal Action Tokenization with Conditional Annealing in Flow Matching
Abstract
Autoregressive Vision-Language-Action (VLA) models offer a scalable path to robot learning, yet existing action tokenizers treat tokenization as a compression problem, producing representations semantically misaligned with the autoregressive backbone. We propose CATok, a causal action tokenizer that reframes tokenization as a causally structured generative process. CATok introduces a conditional annealing mechanism that extracts action tokens by progressively annealing a flow-matching process: each token is conditioned on all preceding tokens and encodes the residual reconstruction signal at a specific noise level, establishing a coarse-to-fine causal token space whose generative semantics are structurally aligned with autoregressive modeling. A token-conditioned flow-matching decoder built on Multimodal Diffusion Transformer (MMDiT) reconstructs continuous action chunks from these discrete tokens. This discrete bottleneck enforces knowledge insulation by design, separating high-level semantic reasoning from low-level motor execution without explicit attention masking. Experiments across three simulation benchmarks and real-world manipulation show that CATok surpasses existing tokenizers in reconstruction fidelity--compression tradeoff and inference efficiency while improving VLA success rate and training efficiency.Our project page is available at https://causalactiontokenizer.github.io.