Gigabytes to Bytes: How Much of a CFD Archive Does a Decision Need?
Abstract
Large power transformers are critical, decades-lived assets of the electrical grid, and their service life is determined largely by cooling. The flow of insulating oil through millimetre-scale winding ducts sets the hot-spot temperature rise, which drives insulation ageing and the peak oil pressure, which the oil circuit must sustain. Each winding design must be evaluated at a given load through a high-fidelity computational fluid dynamics (CFD) solve costing several hours. Neural field surrogates cut this cost by fitting a continuous function to solver-generated data, reproducing the physical fields at high fidelity. However, some of the engineering decisions these simulations support — screening a design of experiments, sizing a pump, checking a temperature rise against a loading limit — consume one or two scalars per case. This paper does not argue against field surrogates; it addresses these use cases, where the decision requires only scalar outputs, not the full field. We investigate how much decision-relevant accuracy is retained when such an archive is reduced to a single 408-byte tabular row per case, a factor of about 1.4 x 10^7 smaller than the median 5.7 GB solver case. On two industrial winding studies with field-surrogate baselines, we evaluate 6 model families per dataset. A conditional rectified flow, evaluated deterministically at a single quadrature node, reaches 5.57% 95th-percentile (q95) relative error on the hydraulic peak pressure, against 4.94% for the field surrogate on identical test cases. On the thermal hot-spot rise, the three best families converge to 1.61-1.62 K mean absolute error (MAE) against 0.603 K for the field surrogate. In both cases, dataset size is decisive: the tabular models approach the field surrogate only once enough cases are available.