ScoutSampler: Learning to Find Causally Critical Frames for Long Video Understanding
Wenjia Jiang ⋅ Chenru Wang ⋅ Yiwei Wang ⋅ yanming yang ⋅ Boyan Han ⋅ Keyu Chen ⋅ Zhou Yang ⋅ Chi Zhang
Abstract
Long video understanding demands that models identify a handful of decisive frames from redundant content under strict context window constraints. Uniform sampling misses temporally critical events, while existing learnable samplers suffer from cold-start instability and unreliable credit assignment under sparse rewards. We propose ScoutSampler, an RL-driven active frame sampling framework that selects frames by their causal relevance to the posed query. A weakly-supervised MIL warm-up first builds a semantic prior through Best-of-N bag mining without frame-level annotation, overcoming cold-start instability; a self-distilled RL stage then resolves credit assignment by coupling a Gumbel-Top-$K$ Student actor with an EMA Teacher that supplies dense per-frame causal scores driving a localized advantage function. Across a wide range of long video benchmarks, ScoutSampler consistently surpasses strong dense sampling baselines with a small fixed frame budget while substantially reducing visual token overhead.
Chat is not available.
Successful Page Load