Watching Itself Watch: Self-Auditing Visual Reliance for Video Reasoning
Rohan Dugad ⋅ Romy Luo ⋅ Kristen Grauman
Abstract
Video reasoning has a distinctive structure: in a chain of thought over video, most tokens follow from prior text, but a small minority hinge on specific moments of the visual input---and which tokens those are varies from one example to the next. Current post-training procedures like GRPO supply supervision that is coarse in exactly this dimension, applying a single scalar reward uniformly across the trajectory while requiring many rollouts and/or costly external judge models. Our key observation is that a model can audit its own visual reliance: at any position, the shift in its next-token distribution when the video is removed from context measures how much that prediction actually used what was seen. Computing this quantity for both a student and a privileged teacher yields a per-token \emph{reliance gap} that localizes positions at which the teacher's predictions depend on the video and the student's do not. We introduce Visually-Aware On-Policy Self-Distillation (VIOS), which turns this self-auditing signal into dense per-token supervision recovered from a single rollout of a single model---no external reward, judge, or larger teacher required. VIOS combines three components: an exponential moving average (EMA) teacher that co-evolves with the student to avoid the ceiling of a frozen rationalizer; a contrastive term that increases the student's sensitivity to the video input; and a reliance gap gate that concentrates supervision on the positions identified. Across 12 benchmarks spanning general, temporal, spatial, knowledge, and long-video understanding, VIOS consistently outperforms its base model and contemporary video-reasoning systems while training in over $9\times$ fewer steps.
Chat is not available.
Successful Page Load