The Adjoint Costate Method: Certifiable Equilibrium Learning under Non-Multiplicative Discounting
Jaemin Seo ⋅ Jaeyong Lee
Abstract
Non-multiplicative discounting generically induces time-inconsistency: the control selected at time $t$ depends on the diagonal of an anchor-indexed costate field, whereas common field-based training objectives average error over the full anchor--time triangle. We show that the analyzed anchored-$L^2$ formulations admit no loss-to-equilibrium-gap certificate, and that an anchor-mixture policy-gradient objective enforces an anchor-averaged rather than diagonal stationarity condition. We propose the Adjoint Costate Method (ACM), which factorizes the linear adjoint equation into discount-free kernels. The discount appears only in a deterministic quadrature read-out, while source and terminal conditions receive positive-measure supervision. For control-independent additive diffusion, verified uniform kernel defects give an a-posteriori equilibrium-gap bound for the policy actually deployed, linear in the defects and without an equilibrium-contraction assumption. Across three stochastic-control benchmarks inside the certified model class and seven solver families, ACM attains the lowest grid-averaged control error.
Chat is not available.
Successful Page Load