Temperature Guidance For Robust Reward Conditioning In Diffusion Planning
Abstract
Diffusion planners enable long-horizon planning through a generative process over full trajectories, mitigating compounding errors of autoregressive methods while handling multimodal futures. Reward conditioning via classifier-free guidance (CFG) yields high-performing plans but is brittle with respect to per-task hyperparameter choices, limiting its scalability. Our analysis reveals that guidance performance hinges on careful adaptation to the data manifold and reward distribution, contributing to CFG's hyperparameter fragility. We propose temperature-guided diffusion planning (TGDP), which adapts CFG to self-calibrate to these characteristics through temperature-conditional sample reweighting during training and adaptive guidance at inference. TGDP mitigates the hyperparameter fragility of CFG and matches state-of-the-art performance across standard benchmarks without per-task tuning, thus enabling more robust and practical diffusion-based planning.