Metamorphic testing of learned AC-OPF evaluation reveals representation dependence
Abstract
Learned AC optimal power flow (AC-OPF) surrogates are usually evaluated using only one grid serialization, so accuracy reported on that encoding says nothing about how the same pipeline behaves on other valid encodings of the same grid. We introduce a metamorphic audit based on thirteen problem-preserving transformations of labels, reference conventions, and electrically equivalent graph structure. We audit 17 model and pipeline configurations, including two trained open-weight checkpoints and 15 random-weight architecture or mechanism controls. Every configuration violates at least six of the thirteen tested relations. We trace discrepancies in the grid foundation model \texttt{GridSFM-Open v1.1} to preprocessing and identify one indexing bug and three noncanonical or numerically unstable feature choices. We also test the transformations as data augmentation for four surrogate models on 500-bus scenarios. Transformed scenarios' error (MAE) decreases for three of four models, and median discrepancy and AC constraint violations decrease for all four. No evaluated prediction meets the study's AC-feasibility threshold. These results show why evaluations of learned physical surrogates should test consistency across equivalent representations alongside accuracy, optimality, and feasibility. For learned power-grid models, including foundation-model checkpoints, these are the criteria by which foundation-model claims should be evaluated under realistic conditions.