$\varphi$TD: Distributional Reinforcement Learning using Characteristic Functions
Thomas Mousseau ⋅ Tyler Kastner ⋅ Davide Baldelli ⋅ Amir-massoud Farahmand
Abstract
Distributional Reinforcement Learning methods aim to learn the entire distribution of returns, yet current algorithmic approaches are limited to a narrow class of distributions. We propose $\varphi$TD, a method to overcome this limitation through computing the loss in the frequency domain. We demonstrate that this further unifies quantile and categorical approaches to distributional RL, as well as enabling the learning of distributions that were previously impossible to capture, such as mixtures of continuous distributions. We empirically study $\varphi$TD across synthetic benchmarks and large-scale deep RL environments, and demonstrate that the flexibility obtained by wider parametric families often leads to improved performance.
Chat is not available.
Successful Page Load