Circular Models Trained with Equilibrium Propagation Unlock Test-Time Computations
Abstract
Predictive coding and equilibrium propagation (EP) perform credit assignment through an energy minimization process. Despite their ability to train models with arbitrary network topologies, they have mostly been used to train feedforward architectures to match the performance of backprop. Here, we use these methods to enable iterative algorithmic reasoning through energy minimization. By closing a feedforward model into a loop, done by identifying its first and last hidden layers, effective recurrent computation emerges from the continuous relaxation of the energy landscape. We show that in predictive coding networks with this circular structure, the bi-level optimization framework of EP provably approximates the gradient update of an implicit layer, allowing it to train a model whose computational depth can increase at test time by relaxing for a longer time for more complex problems. The same construction achieves image classification performance comparable to standard predictive coding models, with substantially fewer parameters. This construction has direct implications for analog hardware, where computational depth could be traded for relaxation time on a fixed-size circuit.