Selecting or Solving? What Best-of-$N$ Verification Delivers in Video Reasoning
Martin Q. Ma ⋅ Yuxiao Qu ⋅ Ruslan Salakhutdinov ⋅ Louis-Philippe Morency ⋅ Paden Tomasello ⋅ Juan Pino
Abstract
Answering a complex question about a long video is a grounding problem: the answer turns on a few seconds of evidence somewhere inside it, and vision--language models (VLMs) remain well below human accuracy at finding them. The standard way to spend extra test-time compute on this kind of video reasoning is best-of-$N$ selection: sample several chain-of-thought answers from a fixed model, then let a second model re-watch the video and pick one. Such pipelines are reported by a single headline number, their margin over a majority vote. That number conflates two effects with opposite deployment consequences: the benefit the second model adds by selecting among candidates, and the benefit it would have supplied by solving the question itself. We separate them, measuring the Selection Increment (SI)---selection accuracy minus the same verifier's accuracy with the candidates withheld---alongside the margin over self-consistency, across five long-video benchmarks, with verifiers spanning weak open models to a frontier closed one. The two prove dissociable, for different reasons. The Selection Increment is largest for the weakest verifiers: a verifier that solves poorly on its own leaves the candidates the most room to help. Beating the majority vote follows the opposite profile: where the vote already sits near its ceiling no verifier we test improves on it; below that ceiling only the strongest do. On Video-Holmes, the two come apart on a single pool of candidates---a weak verifier gains $+19.3$ points from the candidates yet lands $7.1$ points below the vote, while the frontier verifier gains the least of the five we compare and beats the vote. The weakest verifier also shows a failure mode the capable ones resist: on EgoSchema it picks by where an option sits in the list rather than by what the video shows. Read as a merit score, SI puts a verifier that fails to beat a majority vote at the top of its ranking; read as an audit, it says which of the two effects a headline number delivers. Code will be released upon publication.
Chat is not available.
Successful Page Load