DROGO: Default Representation Objective via Graph Optimization in Reinforcement Learning
Abstract
The default representation (DR), originally introduced in neuroscience, and its principal eigenvector have been shown to be effective for a wide variety of applications, including reward shaping, count-based exploration, option discovery, and transfer. However, in prior investigations, the eigenvectors of the DR were computed by first approximating the DR matrix, and then performing an eigendecomposition. This procedure is computationally expensive and does not scale to high-dimensional spaces. In this paper, we propose an objective for approximating the principal eigenvector of the DR using only transitions with a neural network. This objective is inspired by a series of theoretical results and is empirically validated in a number of environments. We then demonstrate its usefulness by applying the learned eigenvectors for reward shaping.