An Empirical Study and Improvement for Speech Emotion Recognition
Multimodal speech emotion recognition aims to detect speakers' emotions from audio and text. Prior works mainly focus on exploiting advanced networks to model and fuse different modality information to facilitate performance, while neglecting the effect of different fusion strategies on emotion recognition. In this work, we consider a simple yet important problem: how to fuse audio and text modality information is more helpful for this multimodal task. Further, we propose a multimodal emotion recognition model improved by perspective loss. Empirical results show our method obtained new state-of-the-art results on the IEMOCAP dataset. The in-depth analysis explains why the improved model can achieve improvements and outperforms baselines.
Code (0)
등록된 구현이 없습니다.
Tasks
Emotion RecognitionMultimodal Emotion RecognitionSpeech Emotion RecognitionSimilar Papers 제목 키워드 기반
Feature Selection Enhancement and Feature Space Visualization for Speech-Based Emotion Recognition
Robust speech emotion recognition relies on the quality of the speech features. We present speech features enhancement strategy that improves speech emotion recognition. We used the INTERSPEECH 2010 challenge feature-set…
Emotion Recognitionfeature selectionSpeech Emotion RecognitionResearch on several key technologies in practical speech emotion recognition
In this dissertation the practical speech emotion recognition technology is studied, including several cognitive related emotion types, namely fidgetiness, confidence and tiredness. The high quality of naturalistic emoti…
ClusteringEmotion RecognitionSpeech Emotion RecognitionHuman Feedback Driven Dynamic Speech Emotion Recognition
This work proposes to explore a new area of dynamic speech emotion recognition. Unlike traditional methods, we assume that each audio track is associated with a sequence of emotions active at different moments in time. T…
Speech Emotion RecognitionSpeech Emotion Recognition Using Speech Feature and Word Embedding
—Emotion recognition can be performed automatically from many modalities. This paper presents a categorical speech emotion recognition using speech features and word embedding. Text features can be combined with speech f…
Emotion RecognitionSpeech Emotion RecognitionLearning Spontaneity to Improve Emotion Recognition In Speech
We investigate the effect and usefulness of spontaneity (i.e. whether a given speech is spontaneous or not) in speech in the context of emotion recognition. We hypothesize that emotional content in speech is interrelated…
Emotion RecognitionSpeech Emotion Recognition