Behavioral and Mechanistic Interpretability Reveal Control-Mechanism Similarity between Gait Policies and Animal Locomotion
Abstract
While deep reinforcement learning (DRL) achieves high performance on gait control tasks, the decision-making process of the learned policy remains a black box. Previous research on behavioral interpretability based on insights from kinematics has analyzed the input and output sequences generated by gait control policies and identified phase structures corresponding to the gait phases of animals. However, because this approach analyzes only inputs and outputs, it cannot reveal the causal mechanisms inside the policy network. Inspired by the fact that animal movements are achieved through combinations of “muscle synergies”—sets of muscles—this study hypothesized that gait control also involves internal, human-interpretable “functional components” similar to muscle synergies, and that gait control tasks are performed through combinations of these components. A functional component is defined as a circuit that controls a small number of joints in a specific direction, and linearly combining these circuits can nearly represent the original policy. In multiple MuJoCo walking benchmarks, policies reconstructed by linearly combining the outputs of each functional component of a pre-trained policy maintained walking performance while retaining at least 89\% of the original total reward and reducing error by up to 66\% compared to policies reconstructed by randomly splitting the output signals. Furthermore, each component controls a subset of joints in a specific direction; by comparing this with the structure of existing gait phases, we can clarify when each component performs control, in what combinations, and in what manner. These results suggest that control policies which seem like black boxes at first glance in fact contain a functional structure similar to muscle synergies in animals. Code and results are available at this link: https://anonymous.4open.science/r/BPA-MPA-F0B6/README.md.