AMAM-128: A Segmentation Benchmark for Annotated Metallic Alloy Microstructures
Abstract
Metallography segmentation is often benchmarked within narrow material and imaging regimes, making it difficult to tell whether progress reflects robust microstructure understanding or dataset-specific tuning. We introduce AMAM-128, a compact stress benchmark designed to make this distinction visible. AMAM-128 contains 128 expert-refined image–mask pairs spanning steel and cast iron, multiple processing conditions, and magnifications from x5 to x50, including non-ideal micrographs with scratches, pores, reflective regions, and ambiguous phase boundaries. We evaluate 45 segmentation approaches, ranging from classical vision and supervised deep networks to foundation and edge-based models. Across this heterogeneous suite, no method comes close to saturating the benchmark, and performance varies substantially across subsets, revealing persistent sensitivity to material, scale, and image regime. AMAM-128 is intended as a controlled comparison resource rather than a large training corpus or an out-of-distribution benchmark. We release the paired dataset, evaluation pipeline, multi-run deep results, and provenance artifacts to support transparent, subset-aware benchmarking of metallography segmentation.