Speech and Text-Based Emotion Recognizer
Affective computing is a field of study that focuses on developing systems and technologies that can understand, interpret, and respond to human emotions. Speech Emotion Recognition (SER), in particular, has got a lot of attention from researchers in the recent past. However, in many cases, the publicly available datasets, used for training and evaluation, are scarce and imbalanced across the emotion labels. In this work, we focused on building a balanced corpus from these publicly available datasets by combining these datasets as well as employing various speech data augmentation techniques. Furthermore, we experimented with different architectures for speech emotion recognition. Our best system, a multi-modal speech, and text-based model, provides a performance of UA(Unweighed Accuracy) + WA (Weighed Accuracy) of 157.57 compared to the baseline algorithm performance of 119.66
Code (0)
등록된 구현이 없습니다.
Tasks
Data AugmentationEmotion RecognitionSpeech Emotion RecognitionSimilar Papers 제목 키워드 기반
Privacy versus Emotion Preservation Trade-offs in Emotion-Preserving Speaker Anonymization
Advances in speech technology now allow unprecedented access to personally identifiable information through speech. To protect such information, the differential privacy field has explored ways to anonymize speech while …
Speaker anonymizationSpeaker VerificationA proposal for Multimodal Emotion Recognition using aural transformers and Action Units on RAVDESS dataset
Emotion recognition is attracting the attention of the research community due to its multiple applications in different fields, such as medicine or autonomous driving. In this paper, we proposed an automatic emotion reco…
Autonomous DrivingEmotion RecognitionFacial Emotion RecognitionMultimodal Emotion Recognition+2Speech Emotion Recognition Based on Self-Attention Weight Correction for Acoustic and Text Features
Speech emotion recognition (SER) is essential for understanding a speaker’s intention. Recently, some groups have attempted to improve SER performance using a bidirectional long short-term memory (BLSTM) to extract featu…
Emotion RecognitionMultimodal Emotion RecognitionSpeech Emotion Recognitionspeech-recognition+1A Speech Recognizer for Frisian/Dutch Council Meetings
We developed a bilingual Frisian/Dutch speech recognizer for council meetings in Fryslân (the Netherlands). During these meetings both Frisian and Dutch are spoken, and code switching between both languages shows up freq…
An Emotion-based Korean Multimodal Empathetic Dialogue System
We propose a Korean multimodal dialogue system targeting emotion-based empathetic dialogues because most research in this field has been conducted in a few languages such as English and Japanese and in certain circumstan…