Behind Looking Right: What LLMs Miss When Reproducing Physics
Abstract
Asked to reproduce a physics paper, an LLM often returns a figure that looks right, yet a match cannot establish independent computation while the answer remains in the input. We mask figures and captions, predefine targets, put the same seven criteria to all runs and have a physicist audit those passing all seven. The solvers fix failures in planning, coding, and execution once reported, but not failures in the governing equations or in figure match. Supplying equations helps recover the model but barely improves figure match; success falls as the path to numerical realisation grows more abstract. Half of the passing papers bypass the prescribed calculation, yet the LLM judge accepts them. Looking right is not being right: reproduction requires auditing the computational route, not only the final artifact.