Metamorphic auditing of learned AC-OPF evaluation reveals representation dependence
Abstract
Evaluation protocols that use one potentially arbitrary encoding of each test case cannot detect whether a model depends on that encoding. Reported capability may therefore reflect both performance on the physical task and dependence on the chosen representation. We demonstrate this problem in learned alternating current (AC) optimal power flow (AC-OPF) surrogates and introduce a metamorphic audit based on thirteen problem-preserving transformations of labels, reference conventions, and electrically equivalent graph structure. We audit 17 model and pipeline configurations, including two trained open-weight checkpoints and 15 random-weight architecture or mechanism components. Every configuration we test violates at least six of the thirteen tested relations. We trace \texttt{GridSFM-Open v1.1}'s discrepancies to preprocessing and identify one indexing bug and three noncanonical or numerically unstable feature choices. We also test the transformations as data augmentation for four surrogate models on 500-bus scenarios. Transformed scenarios' error (MAE) decreases for three of four models, and median discrepancy and AC constraint violations decrease for all four. No evaluated prediction meets the study's AC-feasibility threshold. These results show why evaluations of learned physical surrogates should test consistency across equivalent representations alongside accuracy, optimality, and feasibility.