A systematic evaluation of vision-language models for observational astronomical reasoning tasks
Abstract
Vision-language models (VLMs) are increasingly proposed as general-purpose tools for scientific data interpretation, yet their reliability on real astronomical observations across diverse modalities remains untested. We present AstroVLBench, a benchmark of over 4,100 expert-verified instances across five tasks spanning optical imaging, radio interferometry, multi-wavelength photometry, time-domain light curves, and optical spectroscopy. We systematically evaluate six state-of-the-art frontier models and find that performance is strongly modality-dependent and uncorrelated with general benchmark rankings. Mechanistic ablations reveal three concrete bottlenecks. First, prompts that ground attention in physical mechanisms (why a feature matters) yield more balanced classifications than phenomenological prompts (what to look for). Second, presenting one-dimensional measurements as numerical tables rather than rendered plots improves accuracy by up to 13 percentage points, indicating that plot rendering can obscure physically relevant signal. Third, a qualitative reasoning analysis reveals a pervasive “right-answer-wrong-reason” phenomenon: models routinely reach correct predictions through physically invalid logic, showing that accuracy alone is insufficient for trustworthy scientific deployment. AstroVLBench provides the first systematic, multi-modal baselines for VLMs in observational astronomy and identifies the representation, grounding, and reasoning bottlenecks that must be addressed before VLMs can be trusted as autonomous scientific interpreters.