VideoSTF: Stress-Testing Output Repetition in Video Large Language Models
Yuxin Cao ⋅ Yuxin Cao ⋅ Shangzhi Xu ⋅ Jingling Xue ⋅ Jin Song Dong
Abstract
Video Large Language Models (VideoLLMs) have achieved strong performance on video understanding tasks, yet existing benchmarks evaluate only what models predict, leaving the stability of how they generate largely unexamined. We surface a previously underexplored generation failure of VideoLLMs, defined as output repetition, in which the decoder collapses into self-reinforcing loops of repeated phrases or sentences, and present VideoSTF, a benchmarking framework for systematically measuring, stress-testing, and exploiting this failure mode. VideoSTF formalizes repetition with three complementary $n$-gram-based metrics, ships a standardized testbed of 10,000 diverse videos, and provides a library of controlled temporal stressors. Across 10 advanced VideoLLMs, VideoSTF reveals four key findings: (i) repetition is pervasive and frame-count-invariant, with repetition rates up to 91%; (ii) it spans a severity spectrum from mild redundancy to token-cap loops, and is highly amplified by temporal perturbations; (iii) temporal stressors form a practical black-box attack surface, flipping benign videos into repetitive ones with tens of queries and high attack success rates (up to 98%); and (iv) repetition is a model-level failure decoupled from input redundancy and driven by local temporal disruption, and is robust to decoding, input filtering, prompts, and video distribution shifts, suggesting that mitigation requires architectural or training-level intervention rather than decoding-time fixes. VideoSTF reframes generation stability as a useful evaluation axis for VideoLLMs and provides the tools to study it.
Chat is not available.
Successful Page Load