The VLM as Sensor: Bayesian Active Search for Long Video Understanding
Chong Tang ⋅ Sannara EK ⋅ Dirk Koch ⋅ Robert Mullins ⋅ Alex Weddell ⋅ Jagmohan Chauhan
Abstract
What if the best way to search a video with a vision-language model (VLM) is to stop asking the model to search? Recent approaches give the VLM control over temporal navigation, tying search quality to the model's reasoning ability. We propose the opposite: BeliefSearch treats the VLM as a noisy sensor and delegates navigation to an external Bayesian controller. A single belief over video segments anchors the system: the question sets the prior, each VLM observation updates it, and the same belief decides whether to search, which segment to examine next, and when to stop. Because the belief is explicit, we can trace why each choice was made. The same signal also drives training: each turn earns credit for the uncertainty it reduces, and the final reward is gated by how much the search lowered entropy. This blocks a reward-hacking shortcut where the model otherwise learns to skip search and guess from the preview frames. On long-video benchmarks, our method outperforms all prior search-based methods by up to 8.8 points while using $5$ to $13\times$ fewer frames than recent RL search baselines.
Chat is not available.
Successful Page Load