Tiny Recursive Policy: Recursive Latent Refinement for Visuomotor Control
Abstract
Behaviour cloning has become a powerful approach for learning robotic manipulation policies, and recent advances in action chunking, latent action representations, and generative policies have substantially improved continuous control. However, most policies generate actions using a fixed observation window and do not explicitly retain latent temporal context throughout execution. We introduce Tiny Recursive Policy, a parameter-efficient behaviour-cloning policy that combines persistent temporal state with recursive latent refinement. At each control step, TRP repeatedly updates a pair of high- and low-level latent states using a shared computation block before decoding them into a future action sequence. The resulting latent states are then carried to the next control step, allowing subsequent predictions to incorporate information accumulated throughout the interaction. We evaluate TRP on multi-task, long-horizon, and few-shot manipulation settings in LIBERO against state-of-the-art behaviour cloning baselines, achieving strong performance across evaluation settings while remaining significantly more parameter-efficient.