Hindsight Relabeling is All You Need for Reach-Avoid Learning
Abstract
Offline goal-conditioned reinforcement learning methods have shown promise for reach-avoid tasks, where an agent must reach a target state while avoiding undesirable regions of the state space. Existing approaches typically encode avoid-region information into an augmented state space and cost function, which prevents flexible, dynamic specification of novel avoid-region information at evaluation time. They also rely heavily on meticulously designed reward and cost functions, limiting transferability to novel environments and specifications. We introduce RADT, a decision transformer approach for offline, reward-free, goal-conditioned, avoid-region-conditioned RL. RADT encodes goals and avoid-regions directly as prompt tokens, allowing any number of avoid-regions of arbitrary size to be specified at evaluation time. Using only suboptimal offline trajectories from a random policy, RADT can learn reach-avoid behavior in a completely data-driven manner using our novel avoid-region hindsight relabeling approach, without the need for reward/cost supervision. We benchmark RADT against existing offline goal-conditioned RL models across 17 tasks, environments, and experimental settings. RADT generalizes in a zero-shot manner to out-of-distribution avoid-region sizes and counts, outperforming baselines that require retraining. In one such zero-shot setting, RADT achieves 35.7% improvement in normalized cost over the best retrained baseline while maintaining high goal-reaching success. We also apply RADT to cell reprogramming in biology, demonstrating its versatility.