LOCU: Löwdin-Orthogonalized Constraint Updates for Multi-Constraint Policy Optimization
Abstract
We investigate multi-constraint policy optimization in constrained Markov decision processes (CMDPs), where interactions among constraint gradients often yield ill-conditioned update directions, causing numerically fragile steps and unintended cancellation or redundant overlap in constraint corrections. To address this, we propose LOCU, which decouples constraint interactions by applying symmetric L\"owdin orthogonalization to the natural gradients of the reward and constraints under the Fisher–Rao metric. Combined with Fisher-geometric screening and near-collinear compression, LOCU yields a compact reduced space in which the trust region becomes a Euclidean ball and constraints are handled symmetrically through a low-dimensional active-set solve. For infeasible iterates, LOCU construct a Pareto-descent direction that simultaneously reduces violated constraints while preserving satisfied ones within the same reduced formulation. Experiments show that LOCU remains stable under high constraint coupling, narrow feasible regions, and infeasible starts, and generalizes across different cost critic architectures.