Better Scores Are Not Enough: Evidence-Dependence Evaluation for Video Hallucination
Abstract
Video hallucination refers to outputs unsupported by video evidence, and is commonly associated with weak video grounding and over-dependence on language priors. Yet benchmark accuracy alone does not establish whether correct responses are grounded in video evidence rather than arising from language priors or response bias. When a mitigation method improves a video hallucination benchmark score, does that gain actually depend on video evidence? We propose evidence-dependence evaluation (EDE), a framework that preserves each benchmark’s original evaluation protocol while repeating the evaluation under controlled interventions on visual and temporal evidence. EDE uses four controlled evidence interventions and three complementary diagnostics, requiring neither additional training nor new annotations. Across five benchmarks and two video LLMs, we find that benchmark performance can remain largely unchanged—and in some cases even improve—when visual or temporal evidence is removed or disrupted. On one benchmark, models score higher when the video is removed entirely, and some representation-level mitigation methods exhibit larger condition-matched gains after frame order is destroyed than on the intact video. Complementary diagnostics show that mitigation gains may be driven by error trade-offs or response shifts, and can be sensitive to evaluation format. EDE complements benchmark evaluation by measuring how strongly benchmark performance and reported mitigation gains depend on visual and temporal evidence. Hallucination-reduction claims can thus be evaluated not only by how well a model performs, but also by what video evidence that performance depends on.