CME–SpectrumBench: Can LLMs Analyze Condensed Matter Spectral Data?
Abstract
Large Language Models (LLMs) are presently well assessed by expert–level mathematics and physics benchmarks which focus on theoretical, symbolic reasoning. However, we lack experimental benchmarks that evaluate data analysis in frontier research problems, a capability that future domain–specific foundation models must have. To this end, we present CME–SpectrumBench, an expert–designed benchmark consisting of 190 questions that measures how reliably frontier LLMs analyze research–level spectral scientific data in condensed matter experiments (CME). Using simulated CME spectral data, we prompt models to identify structures from 2D intensity arrays in the face of noise and limited instrumental resolution, a test of dense context reasoning. Models are scored by how accurately their numerical answers approach the ground truth. Under direct query, scores generally pale in comparison to handwritten code due to clear failure modes: models can find simple features but struggle to trace patterns or compute derived quantities, and performance degrades with increased data resolution and noise/convolution levels. LLM performance becomes comparable to handwritten code by allowing tool use (access to a Python interpreter) with/without ReAct prompting, in effect mapping questions onto agentic code generation problems. Our findings show that frontier models can serve as research assistants by interfacing with data through code, but still face difficulties when processing information directly from large datasets — a key barrier future foundation models must overcome.