Towards Benchmarking Autonomous Physics-Analysis Agents
Abstract
Agentic systems are increasingly being explored for automating and expediting high energy particle physics analysis, but an expert is needed to decide whether such an analysis is done correctly and accurately. The challenge of particle physics analyses is that the true result is not known a priori and ``correctness'' is typically determined by extensive tiers of peer review. Further complicating the issue, agents performing published measurements can simply recall or look up the accepted value rather than measure it in data, leading to a scientifically valid result. We present a benchmark for autonomous particle physics analysis targeting an exemplar Higgs analysis in H-->ZZ-->4l channel that addresses the challenges of absent ground truth and expert-in-the-loop evaluation. Our framework includes the construction of blinded pseudo-data with a sealed injected truth, from which the accuracy of a resulting agentic physics analysis can be measured through the extracted physics analysis parameters compared against the sealed truth. This sealed truth combined with a series of robust statistical procedures built on evaluation of a full analysis likelihood allow us to grade the quality of a physics analyses providing an automated framework that quantitatively compares the performance of SOTA agentic systems for physics analysis.