Verification Before Assembly: A Low-Cost Verifier for Generated Scientific Video Frames
Abstract
Generative systems can produce frames for scientific presentation videos that look convincing while silently changing a label, reversing an arrow, or omitting a required component. For scientific video generation, such errors are not merely visual defects: they can change the meaning of the explanation. We use AI-generated scientific visual artifacts as a case study to study verification under imperfect and scarce verification resources, focusing on when inexpensive surrogate evidence is sufficient to act and when uncertain outputs should be escalated to more expensive verification. We build a low-cost verifier that, given a generated frame and its textual visual specification, combines a learned semantic score, zero-shot object-presence score, and a structural score. An XGBoost arbiter integrates these complementary signals to make the final verification decision. The components operate on a shared held-out test set of 300 examples. The semantic judge reaches 75% accuracy, the object judge reaches 87% mismatch recall, and the arbiter retains 84% of matching frames under its recorded match-positive operating point. Rather than treating these numbers as proof of scientific correctness, we use them to study a more basic verification question: what can an inexpensive surrogate verifier establish, and what should happen when its evidence is insufficient? Finally, we propose a claim-level scene-graph formulation that decomposes specifications into inspectable visual predicates to isolate the exact cause of a regeneration decision. The resulting system is intended as a quality-control gate before video assembly, not as an authority on scientific truth.