Beyond FID: Evaluating Synthetic OCT Images for Alzheimer's Disease Screening Using a Multi-Metric Framework
Abstract
Generative AI offers a potential route to addressing severe data scarcity in medical imaging, but whether synthetic images are genuinely useful downstream remains unclear. Conventional metrics such as Fréchet Inception Distance (FID) may indicate improved visual similarity without capturing whether generated samples preserve meaningful biological characteristics of real data, or whether they improve the task they are ultimately intended for. We investigate this limitation for synthetic optical coherence tomography (OCT) generation in Alzheimer's disease (AD) screening, using a standard public dataset of age-matched AD and control patients. We trained three generative approaches independently per class: a variational autoencoder (VAE), a custom GAN, and StyleGAN2-ADA. To ensure a rigorous and clinically meaningful evaluation, we adopted a patient-grouped, nested cross-validation protocol, so that all frames belonging to a given patient were kept strictly within a single fold, preventing any overlap between training and test data. Under this protocol, we evaluated four training configurations: real-only data, and real data augmented separately with each of the three generative methods. Our results in Table 1 show that improved FID does not guarantee classification benefit. StyleGAN2-ADA was the only generator for which a quantitative FID score was computed (FID 99.2 for AD-class outputs), and visual inspection confirmed its outputs were broadly realistic; despite this, real-plus-StyleGAN2-ADA augmentation performed no better than real-only training, and in fact showed the weakest discrimination between classes of all four configurations. VAE and custom GAN augmentation, evaluated qualitatively, performed similarly poorly despite visibly different image quality between the two; full FID computation for all three generators was outside the scope of this abstract and is planned for the full paper. An independent fidelity check further showed that frozen-backbone classifiers strongly misclassified StyleGAN's synthetic images by predicted class, while fully fine-tuned classifiers showed a near-balanced response, suggesting synthetic images carry generic surface-level artifacts rather than genuine class-discriminative anatomical signal. Together, these results demonstrate that visual plausibility and improved FID are not reliable proxies for biomedical usefulness: a synthetic image can look realistic and score well on conventional metrics while still failing to capture the diagnostic signal needed for downstream clinical tasks