FrameSelect: A Unified Library for Video Frame Selection and Evaluation
Abstract
Frame selection---picking a small set of informative frames from a video to feed a vision-language model (VLM)---sits at the front of every video-VLM system, yet the literature is fragmented: most works do not share a candidate pool, encoder, frame budget, or answering VLM, making published gains hard to compare across works. We introduce FrameSelect, to the best of our knowledge the first Python library that treats the selector as a standalone, interchangeable component, cleanly separated from the encoder that computes representations and the VLM that answers. Any registered selector composes with any registered evaluator through a single call, and adding a new selector, dataset, encoder, or VLM backend is an isolated extension; a replay mode further rebuilds any table under a different answering VLM without re-running selection. We release reference implementations of 20 selectors, 15 evaluators spanning answer-level and selection-level metrics, and the first unified benchmark comparing frame selectors under a shared protocol across seven evaluators spanning both evaluation modes. This comparison surfaces findings that existing studies cannot: rankings invert across benchmarks. Our library is available at https://anonymous.4open.science/r/FrameSelect-D2A4.