Towards a Generalization Benchmark for GNN Power-Flow Models: An Out-of-Distribution Evaluation for Transmission Grids
David A. Quispe G.
Abstract
Graph neural networks are increasingly used as fast approximations of AC power flow. They are typically trained and evaluated within a single grid, so it is unclear what their reported performance tells us about a grid they have not seen. We train six GNN architectures with one identical recipe on four transmission grids and evaluate them along a ladder of increasing structural change, from a fixed topology, to random N-1/N-2 line outages on the same grid, to an unseen grid, and to a grid held out of a three grid training mixture. The experiments showed two important findings. First, the ranking of the six architectures survives a harder test on the same grid (Kendall $\tau = 0.62$, exact permutation $p = 0.004$) but not a change of grid ($\tau = 0.07$ and $\tau = 0.00$), so a leaderboard built on one grid does not say which model to pick for another. Second, a prediction can be numerically close to the true state and still not be a physically feasible one. The most accurate model in distribution is off by 42 \% of the load being served, and on an unseen grid the six architectures flag 70 to 79 \% of branches as overloaded while missing 16 to 26 \% of the real overloads. These findings do not rank architectures. They specify what a generalization benchmark could include, and we release the protocol and results as a first step towards a benchmark.
Chat is not available.
Successful Page Load