Learning Reusable Options by Decomposing Neural Policies
Abstract
Options provide a natural form of behavioral abstraction in reinforcement learning, but discovering reusable options from trained neural policies remains difficult. A trained policy may contain many useful behaviors that can be reused in downstream problems, yet the number of candidate subpolicies that a neural network can encode grows exponentially with the number of hidden units. We propose a differentiable approach for extracting reusable options from neural policies. Our key observation is that for piecewise-linear neural networks, selecting a neural subprogram is equivalent to assigning each hidden unit one of three labels: inactive, active, or retained in the computation. This yields a mask space over neural subprograms that can be searched with gradient-based optimization. We further show that reusable options can require default input parameters: neuron masks determine what computation is reused, while input masks determine how that computation is called. Experiments in transfer settings with feedforward and recurrent policies show that the resulting options improve sample efficiency on downstream learning.