Masked Sobolev Training for Feasibility-Reliable Optimization Proxies
Andrew ROSEMBERG ⋅ Joaquim Dias Garcia ⋅ Russell Bent ⋅ Pascal Van Hentenryck
Abstract
Optimization proxies amortize parametric constrained optimization by learning fast approximations of solution maps. Because these maps are smooth only within active-set regimes, the parameter space may split into combinatorially many critical regions, making finite-sample generalization and reliable constraint behavior central challenges. To address this challenge, this paper develops sparse stochastic masked Sobolev training for optimization proxies, augmenting value supervision with randomly selected solver-sensitivity targets rather than dense Jacobian matching. The mask acts as an algorithmic control: it exposes the proxy to local derivative information while avoiding overconstraint from conflicting active constraint set signals. We provide approximation bounds showing that, under coverage and regularity, joint value-and-Jacobian matching improves finite-sample generalization. Experiments characterize when and how this benefit appears. In smooth optimal-control distillation, where feasibility is guaranteed by construction, and only a few active-set regimes arise, Sobolev supervision reaches 95\% of teacher performance about $2.5\times$ faster than value-only training. In AC optimal power flow, masked Sobolev training keeps optimality gaps below 0.22\% across three PGLib networks while reducing the median worst-case constraint violation by up to $4\times$. A mean--variance portfolio study identifies the complementary limitation: when sensitivities are highly local or vanish across active-set regimes, derivative supervision can be less useful than value-only training.
Chat is not available.
Successful Page Load