Tuning the Tuner
Abstract
Properly tuned hyperparameters are critical for reinforcement learning algorithms to perform well, limiting their use in real world applications. AutoRL aims to to address this by nesting the reinforcement learning problem within an outer optimization loop, where a black-box optimizer tunes the reinforcement learning algorithm's hyperparameters leaving the AutoRL algorithm's hyperparameters (hyper-hyperparameters) fixed. We provide the first empirical study of an AutoRL algorithm's sensitivity with respect to its hyper-hyperparameters. Our results demonstrate that SMAC3 with Hyperband, used to tune PPO, is sensitive to its hyper-hyperparameters, albeit less so than PPO is to its hyperparameters, and this reduction comes at a significant computational cost. The per-environment tuned performance gains do not outperform random search, suggesting the reduction in sensitivity may not justify the substantial compute and engineering overhead AutoRL imposes.