Elephant in the Fridge: Constant-Memory Frame Packing for Long Video Understanding
Shuo Gao ⋅ Chenhao Zheng ⋅ Jason Ren
Abstract
Hour-long videos, streaming feeds, and movie-length content are pushing vision-language models (VLMs) past their breaking point—the standard recipe of encoding each frame into visual tokens yields sequences that grow linearly with video length, causing prohibitive memory consumption and inference latency. Existing approaches, such as token reduction and keyframe selection, attempt to shorten the token sequence but leave the $O(n)$ scaling intact, failing to resolve the fundamental bottleneck. We propose \model, \emph{a strikingly simple, training-free frame packing framework that compresses arbitrarily long videos into a fixed-size visual representation at constant memory cost.} \model requires no fine-tuning, no architectural modification, and no auxiliary models—it plugs directly into off-the-shelf VLMs at inference time, yet consistently surpasses more elaborate baselines. Drawing inspiration from FramePack in generation literature, we propose a new content-aware compression strategy that operates under a fixed token budget for visual understanding. Specifically, \model proceeds in two steps: (1) scoring each frame by its semantic similarity to the text prompt and ranking frames in a diversity-driven manner, and (2) packing the selected frames at varying resolutions into a fixed-length context window. This design concentrates the memory budget on frames that are both instruction-relevant and visually diverse. While in theory \model can pack arbitrarily many frames into the fixed budget, we identify an empirical instantiation that generalizes robustly across different settings. Extensive experiments on LongVideoBench, MLVU, and Video-MME show that \model consistently outperforms full-resolution baselines as well as prior token reduction and keyframe selection methods, across diverse model families, parameter scales, and frame budgets. Notably, on Qwen3.5-35B-A3B with a 128-frame budget, \model yields gains of \textbf{+12.9\%, +10.6\%, and +9.7\%} over the official model on the three benchmarks, respectively.
Chat is not available.
Successful Page Load