SCG-HF: Semantic Consistency Grouping for Hierarchical Fusion in Video Emotion Recognition
Abstract
Video Emotion Recognition (VER) aims to identify and understand the emotional states of characters by analyzing visual and auditory information in videos. However, conventional rigid frame sampling strategies for long videos tend to fragment continuous emotional expressions. Moreover, existing fusion methods often fail to properly model temporal dependencies and may suffer from future information leakage. To address these limitations, we propose SCG-HF, a novel framework that preserves semantic consistency through adaptive grouping and hierarchical fusion. Specifically, we introduce Semantic Consistency Grouping (SCG), which dynamically segments videos into semantically coherent units based on feature similarity, thereby maintaining the integrity of emotional patterns. Furthermore, we design a three-level Hierarchical Fusion (HF) architecture to capture emotional dynamics at different temporal granularities: (i) frame-level refinement enhances subtle local emotional cues; (ii) segment-level history-guided fusion models temporal evolution across segments while strictly preventing future information leakage via attention; and (iii) sample-level contrastive alignment synchronizes audio and visual representations in a shared latent space. Extensive experiments on the VideoEmotion-8 and Ekman-6 benchmarks demonstrate that SCG-HF achieves state-of-the-art performance.