Where Does Grounding Fail in Long-Video Question Answering?
Abstract
Grounding in a long-video pipeline is decided before the answering model runs. A frame selector chooses sixteen to sixty-four frames out of several thousand, and evidence it discards cannot support any answer that follows, so the selector fixes what is available to be grounded in. Answer accuracy cannot report on that decision, because a correct answer is consistent with the evidence never arriving. We show this with a controlled counterfactual, substituting a recording that cannot contain the answer while holding the question, the options, the prompt, the decoding settings and the frame budget fixed. Correcting for guessing, substitution removes roughly two thirds of above-chance accuracy on Video-MME and only about half on LongVideoBench, so the extent to which a model needs the intended recording is benchmark-dependent and is not visible in raw accuracy. We therefore evaluate selection directly against annotated evidence, before the answering model produces a token. Published training-free selectors place a minority of the frame budget inside annotated evidence. We present a training-free selector that refines query relevance, uses it to weight frame embeddings, projects onto an effective-rank subspace and selects greedily by residual norm. It exceeds uniform sampling and four published selectors on annotated-evidence precision and scene coverage at all three budgets, reaching 0.415 against 0.204 for the strongest baseline at sixty-four frames on Video-MME. Removing the relevance weighting places the method below query-blind uniform sampling on all six answering models. The results establish improved evidence selection and downstream utility in the tested settings; they do not establish faithful evidence use.