GEMBench: Separating Distribution Matching from Discovery in Crystal Generation
Abstract
Classical generative metrics (Fr\'echet Inception Distance, kernel MMD, and related feature-space distances) score how well a model matches a target distribution. Discovery needs something else: valid samples that go beyond the training support. We call this tension the "generative discovery dilemma": matching and discovery agree on what stability means, but disagree on the pass fraction they reward and on the desired shape of the distribution of crystals that pass those checks, so evaluation must declare the aim. Crystal stability itself has several forms (hull, phonons, finite temperature), yet common yields reduce it to one energy-above-hull bit. GEMBench (Generative Evaluation for Materials) scores the two aims separately as GEM-match and GEM-disc under a shared multi-probe stability panel and shared set-level comparison: GEM-match rewards corpus-like pass rates and shapes, while GEM-disc rewards many distinct novel stables. On ten MP-20 checkpoints, plus a two-model Alex-MP-20 case study, matching, discovery, and CHGNet S.U.N. disagree: which generator leads depends on purpose and on how we define crystal similarity.