ImageNet FID is a Pass Check, Not a Finish Line
Abstract
ImageNet class-conditional generation is a standard first test for new image generation methods. Progress on ImageNet is usually measured with one number: Fréchet Inception Distance (FID), computed by comparing 50K class-balanced generated samples with the ImageNet training set in the Inception-v3 feature space. The FID number was useful when the generation quality was low and the gap between methods was large. However, many recent methods already reach remarkable and close FID numbers, and we argue that the field should stop treating FID as decisive evidence of generation quality. To support this view, we re-evaluate 47 recent ImageNet class-conditional generators while keeping the generated samples fixed and changing only the evaluation protocol. We test multiple variants of FID, including different real reference sets, sample sizes, and feature extractors such as Inception-v3, CLIP, DINOv2, and VGG16. The same generated samples can receive different scores, resulting in varying rankings. Our observation means that a single FID protocol has a risk of being over-optimized and over-interpreted. We do not argue that FID should be removed. Instead, we recomend: keep the conventional ImageNet FID for continuity, but also report the results of some FID varaints to avoid overfitting to one evaluation recipe. If a method reaches comparable and stable ImageNet FID, authors should not be expected to keep chasing the best FID, and the paper should be evaluated based on their technical contributions. After passing the FID test, future work such as text-to-image training and human preference study, is important for the futher development of a generation method.