CADMA: Capacity-Aware Recall Decomposition for Generative Model Assessment
Abstract
Evaluating generative models requires measuring not only sample fidelity but also how well a generated set covers the real data distribution. Recent work has therefore introduced recall-oriented metrics for distribution-level evaluation. However, existing recall metrics often reduce distributional coverage to geometry-based coverage. This simplifying assumption weakens their diagnostic value and interpretability in two ways: these metrics provide only a single recall score that cannot distinguish essential diversity failures such as mode dropping and mode collapse, and geometry-based coverage can provide unreliable estimates by assigning high recall to poorly covered generative distributions. We therefore introduce CADMA, a capacity-aware framework for recall-oriented evaluation and diagnosis. CADMA evaluates distributional coverage by formulating local real--fake allocation as a one-to-one matching problem, separating geometry-based coverage from effective coverage. This yields multi-dimensional recall diagnostics that quantify mode dropping and mode collapse while reducing the overly optimistic estimates produced by geometry-based recall metrics. Through extensive investigations on toy experiments and generative models, we show that CADMA provides more reliable recall estimates and more fine-grained, interpretable diagnostic signals of diversity failures than existing metrics.