CroissantMiner: Automated Extraction and Validation of Croissant Metadata for ML Datasets
Abstract
Croissant has emerged as a community standard for machine-readable dataset metadata, yet populating its fields still requires careful reading of dataset papers and remains a documented pain point for authors. We present CroissantMiner, to our knowledge the first benchmark for end-to-end evaluation of metadata extraction aligned with the full Croissant schema. It comprises 602 ML dataset papers: 102 with human-validated gold annotations (22 annotators, 9,595 ratings) and 500 with LLM-validated silver annotations, covering all 30 core and Responsible AI (RAI) fields. Using the benchmark we evaluate 24 extraction systems spanning frontier models, open-weight models, and four agentic architectures, under a two-tier framework that combines rule-based scoring with an LLM judge selected by human audit (κ = 0.890). Single-pass extraction consistently outperforms every agentic architecture that shares its backbone, at 1.1 to 2.7 times lower cost; the gap is largest on long-form RAI fields, which require synthesizing information scattered across a paper rather than copying it from one location. We release the benchmark, evaluation code, and judge audit.