Learning When to Think: Adaptive Internal Computation for Reinforcement Learning
Abstract
As agents have to solve diverse tasks from grid-worlds to robotics and language modeling, the difficulty of each step varies, yet policies are rigid and spend a fixed amount of compute on each step. A natural approach is to spend compute on the harder states while saving it on easy ones, allowing agents to retain performance while saving valuable compute budgets. To this end, we propose Adaptive Internal Computation (AIC), a lightweight policy architecture that iteratively refines latent states and learns when to halt using a value-of-computation signal.Concretely, AIC repeatedly refines a latent state representation and learns a halting head from a value-of-computation signal, the per-step gain in expected return, allowing the policy to stop refining as soon as further computation no longer improves the action it would take. A central challenge, however, is that matched-budget gains alone cannot establish that compute is being directed to the states that benefit the most from them. We therefore introduce a six-component causal evaluation protocol leveraging a targeted-versus-random ablation that tests whether removing compute from AIC's predicted high-compute states hurts performance more than removing the same amount of compute from randomly chosen states, thereby validating that AIC can successfully identify high-compute states. Across twelve diverse tasks, including grid worlds, simulated robotics experiments, and language modeling, we demonstrate that AIC outperforms fixed-compute policies across a set of five underlying model architectures. In particular, architectures leveraging AIC retain or improve performance while spending about 52\% less compute compared to their respective fixed counterparts.