paper-with-me

Papers

Versatile audio-visual learning for emotion recognition

2023-05-12 · Lucas Goncalves, Seong-Gyun Leem, Wei-Cheng Lin, Berrak Sisman, Carlos Busso

Most current audio-visual emotion recognition models lack the flexibility needed for deployment in practical applications. We envision a multimodal system that works even when only one modality is available and can be implemented interchangeably for either predicting emotional attributes or recognizing categorical emotions. Achieving such flexibility in a multimodal emotion recognition system is difficult due to the inherent challenges in accurately interpreting and integrating varied data sources. It is also a challenge to robustly handle missing or partial information while allowing direct switch between regression or classification tasks. This study proposes a versatile audio-visual learning (VAVL) framework for handling unimodal and multimodal systems for emotion regression or emotion classification tasks. We implement an audio-visual framework that can be trained even when audio and visual paired data is not available for part of the training set (i.e., audio only or only video is present). We achieve this effective representation learning with audio-visual shared layers, residual connections over shared layers, and a unimodal reconstruction task. Our experimental results reveal that our architecture significantly outperforms strong baselines on the CREMA-D, MSP-IMPROV, and CMU-MOSEI corpora. Notably, VAVL attains a new state-of-the-art performance in the emotional attribute prediction task on the MSP-IMPROV corpus.

📄 PDF Abstract BibTeX arXiv:2305.07216

Code (0)

등록된 구현이 없습니다.

Tasks

Arousal EstimationAttributeaudio-visual learningEmotion ClassificationEmotion RecognitionMultimodal Emotion RecognitionregressionRepresentation LearningSpeech Emotion RecognitionVideo Emotion Recognition

Similar Papers 제목 키워드 기반

A vector quantized masked autoencoder for audiovisual speech emotion recognition

2023-05-05 · Samir Sadok, Simon Leglaive, Renaud Séguier

An important challenge in emotion recognition is to develop methods that can leverage unlabeled training data. In this paper, we propose the VQ-MAE-AV model, a self-supervised multimodal model that leverages masked autoe…

Contrastive LearningEmotion RecognitionRepresentation LearningSelf-Supervised Learning+1

Audio Visual Emotion Recognition with Temporal Alignment and Perception Attention

2016-03-28 · Linlin Chao, Jian-Hua Tao, Minghao Yang, Ya Li 외

This paper focuses on two key problems for audio-visual emotion recognition in the video. One is the audio and visual streams temporal alignment for feature level fusion. The other one is locating and re-weighting the pe…

ClassificationEmotion RecognitionGeneral Classification

Tailor Versatile Multi-modal Learning for Multi-label Emotion Recognition

2022-01-15 · Yi Zhang, Mingyuan Chen, Jundong Shen, Chongjun Wang

Multi-modal Multi-label Emotion Recognition (MMER) aims to identify various human emotions from heterogeneous visual, audio and text modalities. Previous methods mainly focus on projecting multiple modalities into a comm…

DecoderDiversityEmotion Recognition

Emotion Recognition System from Speech and Visual Information based on Convolutional Neural Networks

2020-02-29 · Nicolae-Catalin Ristea, Liviu Cristian Dutu, Anamaria Radoi

Emotion recognition has become an important field of research in the human-computer interactions domain. The latest advancements in the field show that combining visual with audio information lead to better results if co…

Emotion Recognition

An End-to-End Visual-Audio Attention Network for Emotion Recognition in User-Generated Videos

2020-02-12 · Sicheng Zhao, Yunsheng Ma, Yang Gu, Jufeng Yang 외

Emotion recognition in user-generated videos plays an important role in human-centered computing. Existing methods mainly employ traditional two-stage shallow pipeline, i.e. extracting visual and/or audio features and tr…

Emotion RecognitionVideo Emotion Recognition