Reusing Neural Speech Representations for Auditory Emotion Recognition
Acoustic emotion recognition aims to categorize the affective state of the speaker and is still a difficult task for machine learning models. The difficulties come from the scarcity of training data, general subjectivity in emotion perception resulting in low annotator agreement, and the uncertainty about which features are the most relevant and robust ones for classification. In this paper, we will tackle the latter problem. Inspired by the recent success of transfer learning methods we propose a set of architectures which utilize neural representations inferred by training on large speech databases for the acoustic emotion recognition task. Our experiments on the IEMOCAP dataset show ~10% relative improvements in the accuracy and F1-score over the baseline recurrent neural network which is trained end-to-end for emotion recognition.
Code (0)
등록된 구현이 없습니다.
Tasks
Emotion RecognitionGeneral ClassificationTransfer LearningSimilar Papers 제목 키워드 기반
Research on several key technologies in practical speech emotion recognition
In this dissertation the practical speech emotion recognition technology is studied, including several cognitive related emotion types, namely fidgetiness, confidence and tiredness. The high quality of naturalistic emoti…
ClusteringEmotion RecognitionSpeech Emotion RecognitionMFHCA: Enhancing Speech Emotion Recognition Via Multi-Spatial Fusion and Hierarchical Cooperative Attention
Speech emotion recognition is crucial in human-computer interaction, but extracting and using emotional cues from audio poses challenges. This paper introduces MFHCA, a novel method for Speech Emotion Recognition using M…
Emotion RecognitionSpeech Emotion RecognitionJointly Learning Visual and Auditory Speech Representations from Raw Data
We present RAVEn, a self-supervised multi-modal approach to jointly learn visual and auditory speech representations. Our pre-training objective involves encoding masked inputs, and then predicting contextualised targets…
Audio-Visual Speech RecognitionLipreadingspeech-recognitionSpeech Recognition+1Toward Efficient Speech Emotion Recognition via Spectral Learning and Attention
Speech Emotion Recognition (SER) traditionally relies on auditory data analysis for emotion classification. Several studies have adopted different methods for SER. However, existing SER methods often struggle to capture …
Speech Emotion RecognitionEmotion ClassificationData AugmentationEnd-to-End Multimodal Emotion Recognition using Deep Neural Networks
Automatic affect recognition is a challenging task due to the various modalities emotions can be expressed with. Applications can be found in many domains including multimedia retrieval and human computer interaction. In…
Emotion RecognitionMultimodal Emotion RecognitionRetrieval