ExoBench: Multi-Expert Evaluation of Agentic End-to-End Exoplanet Transit Analysis and Scientific Report Generation
Abstract
The transit method for exoplanets is one of the primary techniques for detecting and characterizing planets beyond the Solar System, yet its analysis pipeline remains tedious. In this paper, we present ExoBench, a benchmark containing 20 unique exoplanet observations that decomposes the end-to-end transit-analysis pipeline into four subtasks: target-star identification, EXOTIC pipeline execution, transit parameter extraction, and scientific report generation. ExoBench combines automated metrics with expert evaluation of generated scientific reports, enabling both automated verification of outputs and expert validation of the final scientific interpretation. We evaluate six ExoAgent configurations across ExoBench and find that current tool-use LLM agents can complete much of the pipeline autonomously, while still requiring expert verification for reliable scientific use.