Multi-Granularity Network with Modal Attention for Dense Affective Understanding
Video affective understanding, which aims to predict the evoked expressions by the video content, is desired for video creation and recommendation. In the recent EEV challenge, a dense affective understanding task is proposed and requires frame-level affective prediction. In this paper, we propose a multi-granularity network with modal attention (MGN-MA), which employs multi-granularity features for better description of the target frame. Specifically, the multi-granularity features could be divided into frame-level, clips-level and video-level features, which corresponds to visual-salient content, semantic-context and video theme information. Then the modal attention fusion module is designed to fuse the multi-granularity features and emphasize more affection-relevant modals. Finally, the fused feature is fed into a Mixtures Of Experts (MOE) classifier to predict the expressions. Further employing model-ensemble post-processing, the proposed method achieves the correlation score of 0.02292 in the EEV challenge.
Code (0)
등록된 구현이 없습니다.
Similar Papers 제목 키워드 기반
Multi-Granularity Sentiment Integration for LLM-Based Multimodal Sentiment Analysis
Multimodal sentiment analysis (MSA) aims to predict sentiment polarity and intensity from heterogeneous inputs such as text, audio, and vision. While large language models (LLMs) offer strong semantic priors for MSA, eff…
Multimodal Sentiment AnalysisMDAN: Multi-level Dependent Attention Network for Visual Emotion Analysis
Visual Emotion Analysis (VEA) is attracting increasing attention. One of the biggest challenges of VEA is to bridge the affective gap between visual clues in a picture and the emotion expressed by the picture. As the gra…
Emotion RecognitionRecent Trends of Multimodal Affective Computing: A Survey from NLP Perspective
Multimodal affective computing (MAC) has garnered increasing attention due to its broad applications in analyzing human behaviors and intentions, especially in text-dominated multimodal affective computing field. This su…
Aspect-Based Sentiment AnalysisEmotion RecognitionEmotion Recognition in ConversationMultimodal Emotion Recognition+3Hierarchical Granularity Alignment and State Space Modeling for Robust Multimodal AU Detection in the Wild
Facial Action Unit (AU) detection in in-the-wild environments remains a formidable challenge due to severe spatial-temporal heterogeneity, unconstrained poses, and complex audio-visual dependencies. While recent multimod…
AffectVerse: Emotional World Models for Multimodal Affective Computing
Humans infer emotions by integrating observed multimodal cues with expectations about how affective states may unfold. Existing multimodal large language models (MLLMs), however, often treat emotion recognition as static…
Emotion Recognition