From Label Priors to Task Evidence: Long-Video Frame Selection via Bayesian GRPO
Abstract
Multimodal Large Language Models (MLLMs) face significant scalability challenges in long-video understanding due to quadratic attention complexity and context window constraints. Existing retrieval-based approaches typically using supervised learning on static keyframe annotations. However, these label priors often fail to align with the actual information required by the downstream MLLM for complex reasoning, creating a fundamental misalignment between the training target and inference needs. To bridge this gap, we propose a probabilistic framework that reformulates frame selection as a Bayesian inference process. We first initialize a policy using supervised learning to capture general saliency. Then, we treat the frozen MLLM as a task environment and employ its negative log-likelihood (NLL) as dense reward evidence to update the policy via Bayesian Group Relative Policy Optimization (GRPO). This process explicitly transits the selection principle from label priors to task evidence. Extensive experiments on Video-MME, MLVU, and LongVideoBench demonstrate that our approach matches the accuracy of heavy MLLM scorers while maintaining the efficiency of lightweight models.