SCOPE: Sparse-Context On-Policy Distillation for Efficient Vision Language Models
Abstract
Vision-language models (VLMs) process images and videos as long sequences of visual tokens, making inference computationally expensive. Training-free token pruning can substantially reduce this cost by retaining only a compact subset of visual tokens, but performance often sharply degrades at aggressive pruning settings. This degradation is commonly attributed to the irreversible loss of task-relevant visual information. In this work, we revisit this assumption. Through a fixed-context recoverability analysis, we uncover a striking discrepancy: while greedy pass@1 accuracy drops substantially under aggressive pruning, pass@k recovers a significant fraction of the lost performance under repeated sampling. This suggests that, for many examples, pruning does not eliminate all task-relevant information; rather, the resulting sparse visual representation becomes difficult for the language model to reliably interpret and exploit. We characterize this phenomenon as a representation--utilization gap. Motivated by this observation, we introduce SCOPE, a lightweight post-training framework based on on-policy self-distillation. The student first generates trajectories conditioned on the pruned visual context, while a teacher---the same model with access to the corresponding full visual context as privileged information---provides dense token-level supervision along the student's on-policy prefixes. But not all failures are caused by a lack of utilization and aggressive pruning will sometimes remove information required to solve the task. We therefore introduce a selective distillation mechanism that identifies cases for which the pruned context remains sufficiently informative, avoiding misleading supervision. Trained with only 10% of the visual tokens in the student context, SCOPE improves the eight-benchmark normalized aggregate from 85.03 to 90.98 at the same inference budget---a 7.0% relative gain over the pruned vanilla baseline, recovering approximately 40% of the performance lost to pruning. Our results show that improving a VLM's ability to use sparse visual representations can complement better pruning itself, enabling more efficient deployment without architectural modifications.