A Transfer Learning Method for Speech Emotion Recognition from Automatic Speech Recognition
This paper presents a transfer learning method in speech emotion recognition based on a Time-Delay Neural Network (TDNN) architecture. A major challenge in the current speech-based emotion detection research is data scarcity. The proposed method resolves this problem by applying transfer learning techniques in order to leverage data from the automatic speech recognition (ASR) task for which ample data is available. Our experiments also show the advantage of speaker-class adaptation modeling techniques by adopting identity-vector (i-vector) based features in addition to standard Mel-Frequency Cepstral Coefficient (MFCC) features.[1] We show the transfer learning models significantly outperform the other methods without pretraining on ASR. The experiments performed on the publicly available IEMOCAP dataset which provides 12 hours of motional speech data. The transfer learning was initialized by using the Ted-Lium v.2 speech dataset providing 207 hours of audio with the corresponding transcripts. We achieve the highest significantly higher accuracy when compared to state-of-the-art, using five-fold cross validation. Using only speech, we obtain an accuracy 71.7% for anger, excitement, sadness, and neutrality emotion content.
Code (0)
등록된 구현이 없습니다.
Tasks
Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Emotion RecognitionSpeech Emotion Recognitionspeech-recognitionSpeech RecognitionTransfer LearningSimilar Papers 제목 키워드 기반
ASR-based Features for Emotion Recognition: A Transfer Learning Approach
During the last decade, the applications of signal processing have drastically improved with deep learning. However areas of affecting computing such as emotional speech synthesis or emotion recognition from spoken langu…
Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Emotional Speech SynthesisEmotion Recognition+4Towards Multimodal Emotion Recognition in German Speech Events in Cars using Transfer Learning
The recognition of emotions by humans is a complex process which considers multiple interacting signals such as facial expressions and both prosody and semantic content of utterances. Commonly, research on automatic reco…
Emotion RecognitionMultimodal Emotion RecognitionTransfer LearningTransfer Learning for Improving Speech Emotion Classification Accuracy
The majority of existing speech emotion recognition research focuses on automatic emotion detection using training and testing data from same corpus collected under the same conditions. The performance of such systems ha…
ClassificationCross-corpusEmotion ClassificationEmotion Recognition+3DeepEmoNet: Building Machine Learning Models for Automatic Emotion Recognition in Human Speeches
Speech emotion recognition (SER) has been a challenging problem in spoken language processing research, because it is unclear how human emotions are connected to various components of sounds such as pitch, loudness, and …
Speech Emotion RecognitionTransfer LearningData AugmentationEmotion Recognition in Speech using Cross-Modal Transfer in the Wild
Obtaining large, human labelled speech datasets to train models for emotion recognition is a notoriously challenging task, hindered by annotation cost and label ambiguity. In this work, we consider the task of learning e…
Emotion RecognitionFacial Emotion RecognitionFacial Expression Recognition (FER)Speech Emotion Recognition