Most claims cannot be checked: coverage limits on instrument-grounded verification of LLM scientific explanations
Abstract
Language models write fluent causal accounts of scientific observations, and a sound account is hard to tell from a merely plausible one. One response is a deterministic verifier that scores each statement against independent measurements. We built one and ran it over 1,781 statements that three models wrote about 50 air quality anomaly events in Houston, Texas, detected in 405,432 observations from eight instrument networks. The verifier assigns each statement to one of ten types and checks it only against measurement channels that can physically bear on that type. When no channel can speak, it abstains. It abstains a lot. Of 1,781 statements, 860 never reach the channel stage: their terms, quantities or named sources cannot be reconciled with the record the model was shown. Nine in ten of those name the right sources anyway, so this is not careless citation so much as a mismatch about what the record says. Of the 921 that do reach it, 407 have no independent channel able to weigh in. Nothing in the corpus was checkable against more than two channels. Abstention is also not a fixed property of the sensor network: on the same events, the same stored evidence and the same rules, it ranges from 58.2% for one model to 83.0% for another. What the generator chooses to say determines most of it. We argue that coverage, not accuracy, is the binding constraint on surrogate verifiers, and that it belongs in the results table rather than the limitations section. We also describe a preregistered test of whether the verifier's verdicts track expert judgment. Those labels are held out and have not been collected. We claim no agreement result here.