Spatial-temporal Attributes Enhanced Prompt Weighting and Fusion for Video Recognition
Weilin Qiu ⋅ Kaiyan Cao ⋅ Jinhua Ma
Abstract
Adapting large-scale vision-language models (VLMs) such as CLIP to video understanding has drawn increasing attention. Prompt tuning offers an efficient and promising way to adapt VLMs to this task. While existing methods propose to enrich textual information over simple textual templates via large language models (LLMs), the generated descriptions are often coarse and may lead to semantic misalignment with the specific video. To address this issue, we propose $\textbf{S}$patial-$\textbf{T}$emporal $\textbf{A}$ttributes enhanced prompt $\textbf{W}$eighting and $\textbf{F}$usion (ST-AWF), a novel parameter-efficient fine-tuning framework for various video recognition tasks. In our method, LLM is leveraged to generate spatial and temporal attributes for each action category, enriching text prompts with informative descriptions. To mitigate attribute irrelevance in LLM outputs, we design a vision-guided prompt weighting mechanism that dynamically evaluates the relevance of each attribute embedding based on its similarity to video features, thereby emphasizing discriminative cues. Furthermore, a dual-attention cross-modal fusion module is proposed to align weighted textual prompts of spatial–temporal attributes with video features for fine-grained cross-modal feature enhancement. Extensive experiments demonstrate that ST-AWF achieves state-of-the-art performance on multiple video recognition benchmarks, consistently improving both discriminability and generalization ability to unseen classes, while maintaining high efficiency with a small number of trainable parameters. Code is available at https://anonymous.4open.science/r/ST-AWF-code/
Chat is not available.
Successful Page Load