What Do RL Spacecraft Guidance Policies Learn? Symbolic Distillation against a Known Control Law
Joun Won
Abstract
Reinforcement learning (RL) produces effective low-thrust spacecraft guidance policies, but what they learn remains opaque. We train five independently seeded SAC policies on a planar orbit-raising task, distill each into closed-form laws via symbolic regression, and compare the recovered structure against Q-law. The raw networks learn qualitatively different switching behaviors. Their distilled expressions nonetheless collapse onto a common two-channel template in 10 of 15 distillations, through the equinoctial identities alone. The remaining five match on one channel only, but the tangential gain varies by only $7.8\%$ across all fifteen. The template is the small-eccentricity linearization of Q-law's steering. A two-line law built from the parameter medians, with no continuous-parameter tuning, succeeds on every evaluation episode at lower fuel cost than any teacher (paired $-7.2\%$), and matches Q-law's cost ($-1.6\%$ against the untuned baseline, $-0.6\%$ once Q-law's weights are tuned). Agreement with Lyapunov-gradient steering is in part built into the objective, since the reward is itself a Lyapunov-type tracking error. What the objective does not supply is the convergence onto this particular parameterization with a seed-stable gain. Only 6 of 15 tangential expressions are exactly invariant to the periapsis angle, a symmetry Q-law respects, so the template holds at fixed periapsis rather than globally. Code, data, and distilled equations: https://anonymous.4open.science/r/glassbox-guidance1-360C/
Chat is not available.
Successful Page Load