Do FID Evaluations Bias Generative Models Away from Statistical Relevance?
Saumya Goyal ⋅ Dhruv Garg ⋅ Divyan Goyal ⋅ Barnabas Poczos
Abstract
Generative modeling research has increasingly converged on the Fréchet Inception Distance (FID) as its primary evaluation metric, largely due to its computational and statistical convenience relative to statistical metrics such as total variation (TV) and Wasserstein-1 (W1) distance. In this work we aim to study whether an emphasis on FID has driven the field towards models that appear to perform well without a corresponding statistical improvement. We formalize model performance on a metric via its sample complexity: the rate at which error on that metric decreases as a function of the number of training samples $n$. Motivated by the sparse-prior framework of Goyal and Póczos, [2026], which predicts efficient, near-parametric sample complexity of $\tilde{O}(n^{-1/2})$ under a sparsity assumption of realistic priors over distributions of interest, we first empirically verify the validity of their assumptions on priors over the Meta-Dataset and ImageNet. We then measure the sample complexity of two popular generative approaches, DDPM and a flow-matching model, on CelebA and CIFAR-10 under FID, TV, and W1. We find slow sample complexity rates on FID that do not translate to rates on TV and W1, and further that improvements on one metric do not correspond to improvements on other metrics. These results motivate further investigation into the trust we place on FID evaluations, and into discovering metrics that are statistically relevant while being computationally efficient and closer to human perception.
Chat is not available.
Successful Page Load