paper-with-me

홈 › Papers

Temporal aggregation of audio-visual modalities for emotion recognition

2020-07-08 · Andreea Birhala, Catalin Nicolae Ristea, Anamaria Radoi, Liviu Cristian Dutu

Emotion recognition has a pivotal role in affective computing and in human-computer interaction. The current technological developments lead to increased possibilities of collecting data about the emotional state of a person. In general, human perception regarding the emotion transmitted by a subject is based on vocal and visual information collected in the first seconds of interaction with the subject. As a consequence, the integration of verbal (i.e., speech) and non-verbal (i.e., image) information seems to be the preferred choice in most of the current approaches towards emotion recognition. In this paper, we propose a multimodal fusion technique for emotion recognition based on combining audio-visual modalities from a temporal window with different temporal offsets for each modality. We show that our proposed method outperforms other methods from the literature and human accuracy rating. The experiments are conducted over the open-access multimodal dataset CREMA-D.

📄 PDF Abstract BibTeX arXiv:2007.04364

Code (0)

등록된 구현이 없습니다.

Tasks

Emotion Recognition

Similar Papers 제목 키워드 기반

Leveraging Recent Advances in Deep Learning for Audio-Visual Emotion Recognition

2021-03-16 · Liam Schoneveld, Alice Othmani, Hazem Abdelkawy

Emotional expressions are the behaviors that communicate our emotional state or attitude to others. They are expressed through verbal and non-verbal communication. Complex human behavior can be understood by studying phy…

Deep LearningEmotion RecognitionFacial Expression Recognition (FER)Knowledge Distillation

Enriching Multimodal Sentiment Analysis through Textual Emotional Descriptions of Visual-Audio Content

2024-12-12 · Sheng Wu, Xiaobao Wang, Longbiao Wang, Dongxiao He 외

Multimodal Sentiment Analysis (MSA) stands as a critical research frontier, seeking to comprehensively unravel human emotions by amalgamating text, audio, and visual data. Yet, discerning subtle emotional nuances within …

Multimodal Sentiment AnalysisSentiment Analysis

Recursive Joint Cross-Modal Attention for Multimodal Fusion in Dimensional Emotion Recognition

2024-03-20 · R. Gnana Praveen, Jahangir Alam

Though multimodal emotion recognition has achieved significant progress over recent years, the potential of rich synergic relationships across the modalities is not fully exploited. In this paper, we introduce Recursive …

Emotion RecognitionMultimodal Emotion Recognition

EPIC-Fusion: Audio-Visual Temporal Binding for Egocentric Action Recognition

2019-08-22 · ICCV 2019 10 · Evangelos Kazakos, Arsha Nagrani, Andrew Zisserman, Dima Damen

We focus on multi-modal fusion for egocentric action recognition, and propose a novel architecture for multi-modal temporal-binding, i.e. the combination of modalities within a range of temporal offsets. We train the arc…

Action RecognitionEgocentric Activity Recognition

Multi-modal Aggregation for Video Classification

2017-10-27 · Chen Chen, Xiaowei Zhao, Yang Liu

In this paper, we present a solution to Large-Scale Video Classification Challenge (LSVC2017) [1] that ranked the 1st place. We focused on a variety of modalities that cover visual, motion and audio. Also, we visualized …

ClassificationGeneral ClassificationVideo Classification