End-to-End Speech Emotion Recognition: Challenges of Real-Life Emergency Call Centers Data Recordings
Recognizing a speaker's emotion from their speech can be a key element in emergency call centers. End-to-end deep learning systems for speech emotion recognition now achieve equivalent or even better results than conventional machine learning approaches. In this paper, in order to validate the performance of our neural network architecture for emotion recognition from speech, we first trained and tested it on the widely used corpus accessible by the community, IEMOCAP. We then used the same architecture as the real life corpus, CEMO, composed of 440 dialogs (2h16m) from 485 speakers. The most frequent emotions expressed by callers in these real life emergency dialogues are fear, anger and positive emotions such as relief. In the IEMOCAP general topic conversations, the most frequent emotions are sadness, anger and happiness. Using the same end-to-end deep learning architecture, an Unweighted Accuracy Recall (UA) of 63% is obtained on IEMOCAP and a UA of 45.6% on CEMO, each with 4 classes. Using only 2 classes (Anger, Neutral), the results for CEMO are 76.9% UA compared to 81.1% UA for IEMOCAP. We expect that these encouraging results with CEMO can be improved by combining the audio channel with the linguistic channel. Real-life emotions are clearly more complex than acted ones, mainly due to the large diversity of emotional expressions of speakers. Index Terms-emotion detection, end-to-end deep learning architecture, call center, real-life database, complex emotions.
Code (0)
등록된 구현이 없습니다.
Tasks
Deep LearningEmotion RecognitionSpeech Emotion RecognitionSimilar Papers 제목 키워드 기반
End-to-End Continuous Speech Emotion Recognition in Real-life Customer Service Call Center Conversations
Speech Emotion recognition (SER) in call center conversations has emerged as a valuable tool for assessing the quality of interactions between clients and agents. In contrast to controlled laboratory environments, real-l…
Emotion RecognitionSpeech Emotion RecognitionIndian EmoSpeech Command Dataset: A dataset for emotion based speech recognition in the wild
Speech emotion analysis is an important task which further enables several application use cases. The non-verbal sounds within speech utterances also play a pivotal role in emotion analysis in speech. Due to the widespre…
Emotion RecognitionKeyword Spottingspeech-recognitionSpeech RecognitionSpeech Emotion Diarization: Which Emotion Appears When?
Speech Emotion Recognition (SER) typically relies on utterance-level solutions. However, emotions conveyed through speech should be considered as discrete speech events with definite temporal boundaries, rather than attr…
Emotion Recognitionspeaker-diarizationSpeaker DiarizationSpeech Emotion RecognitionMulti-Microphone Speech Emotion Recognition using the Hierarchical Token-semantic Audio Transformer Architecture
The performance of most emotion recognition systems degrades in real-life situations ('in the wild' scenarios) where the audio is contaminated by reverberation. Our study explores new methods to alleviate the performance…
Emotion ClassificationEmotion RecognitionSpeech Emotion RecognitionExploring Attention Mechanisms for Multimodal Emotion Recognition in an Emergency Call Center Corpus
The emotion detection technology to enhance human decision-making is an important research issue for real-world applications, but real-life emotion datasets are relatively rare and small. The experiments conducted in thi…
Decision MakingEmotion RecognitionMultimodal Emotion RecognitionSpeech Emotion Recognition