paper-with-me

Papers

Video Affective Effects Prediction with Multi-modal Fusion and Shot-Long Temporal Context

2019-09-01 · Jie Zhang, Yin Zhao, Longjun Cai, Chaoping Tu, Wu Wei

Predicting the emotional impact of videos using machine learning is a challenging task considering the varieties of modalities, the complicated temporal contex of the video as well as the time dependency of the emotional states. Feature extraction, multi-modal fusion and temporal context fusion are crucial stages for predicting valence and arousal values in the emotional impact, but have not been successfully exploited. In this paper, we propose a comprehensive framework with novel designs of modal structure and multi-modal fusion strategy. We select the most suitable modalities for valence and arousal tasks respectively and each modal feature is extracted using the modality-specific pre-trained deep model on large generic dataset. Two-time-scale structures, one for the intra-clip and the other for the inter-clip, are proposed to capture the temporal dependency of video content and emotion states. To combine the complementary information from multiple modalities, an effective and efficient residual-based progressive training strategy is proposed. Each modality is step-wisely combined into the multi-modal model, responsible for completing the missing parts of features. With all those improvements above, our proposed prediction framework achieves better performance on the LIRIS-ACCEDE dataset with a large margin compared to the state-of-the-art.

📄 PDF Abstract BibTeX arXiv:1909.01763

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Multi-Granularity Network with Modal Attention for Dense Affective Understanding

2021-06-18 · Baoming Yan, Lin Wang, Ke Gao, Bo Gao 외

Video affective understanding, which aims to predict the evoked expressions by the video content, is desired for video creation and recommendation. In the recent EEV challenge, a dense affective understanding task is pro…

Facial Expression and Peripheral Physiology Fusion to Decode Individualized Affective Experience

2018-11-18 · Yu Yin, Mohsen Nabian, Miolin Fan, Chun-An Chou 외

In this paper, we present a multimodal approach to simultaneously analyze facial movements and several peripheral physiological signals to decode individualized affective experiences under positive and negative emotional…

Prediction

AffectVerse: Emotional World Models for Multimodal Affective Computing

2026-05-19 · Bo Zhao, Fanghua Ye, Yixin Ji, Sicheng Zhao 외 arxiv

Humans infer emotions by integrating observed multimodal cues with expectations about how affective states may unfold. Existing multimodal large language models (MLLMs), however, often treat emotion recognition as static…

Emotion Recognition

MART: Masked Affective RepresenTation Learning via Masked Temporal Distribution Distillation

2024-01-01 · CVPR 2024 1 · Zhicheng Zhang, Pancheng Zhao, Eunil Park, Jufeng Yang

Limited training data is a long-standing problem for video emotion analysis (VEA). Existing works leverage the power of large-scale image datasets for transferring while failing to extract the temporal correlation of…

Emotion RecognitionMultimodal Emotion RecognitionMultimodal Sentiment AnalysisRepresentation Learning+2

Multimodal Deep Models for Predicting Affective Responses Evoked by Movies

2019-09-16 · Ha Thi Phuong Thao, Dorien Herremans, Gemma Roig

The goal of this study is to develop and analyze multimodal models for predicting experienced affective responses of viewers watching movie clips. We develop hybrid multimodal prediction models based on both the video an…

Optical Flow Estimation