CroissantMiner: Automated Extraction and Validation of Croissant Metadata for ML Datasets
Abstract
Croissant has emerged as a standard for machine-readable dataset metadata, yet populating its fields remains labor-intensive and requires careful reading of accompanying dataset documentation. The benchmark comprises 602 papers, including 102 with human-validated gold annotations (9,595 ratings) and 500 with LLM-validated silver annotations, covering the full Croissant schema with both core and Responsible AI (RAI) fields. To our knowledge, this is the first benchmark enabling end-to-end evaluation of metadata extraction aligned with a community-standard schema. We evaluate 20 extraction systems, including frontier and open-weight models as well as four agentic architectures, using a two-tier evaluation framework that combines rule-based scoring with an LLM judge selected via human audit. Our results show that single-pass extraction with Claude Sonnet 4.6 and a canonical prompt is the strongest configuration (composite 0.704). Across all backbones, agentic architectures consistently underperform their single-pass counterparts, indicating that task decomposition degrades performance for schema-complete metadata extraction. This effect is most prevalent on long-form RAI fields, where accurate extraction requires integrating information across multiple sections. Claude Sonnet 4.5, used to seed the gold pre-fills, is excluded from the headline ranking to avoid circularity. We release the benchmark, evaluation code, and judge audit anonymously for review at https://anonymous.4open.science/r/croissantminer-F87C, with a live Hugging Face Space demo to be provided at camera-ready.