A Change of Heart: Improving Speech Emotion Recognition through Speech-to-Text Modality Conversion
Speech Emotion Recognition (SER) is a challenging task. In this paper, we introduce a modality conversion concept aimed at enhancing emotion recognition performance on the MELD dataset. We assess our approach through two experiments: first, a method named Modality-Conversion that employs automatic speech recognition (ASR) systems, followed by a text classifier; second, we assume perfect ASR output and investigate the impact of modality conversion on SER, this method is called Modality-Conversion++. Our findings indicate that the first method yields substantial results, while the second method outperforms state-of-the-art (SOTA) speech-based approaches in terms of SER weighted-F1 (WF1) score on the MELD dataset. This research highlights the potential of modality conversion for tasks that can be conducted in alternative modalities.
Code (1)
Tasks
Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Emotion RecognitionSpeech Emotion Recognitionspeech-recognitionSpeech RecognitionSpeech-to-TextSimilar Papers 제목 키워드 기반
Speech Signal Analysis for the Estimation of Heart Rates Under Different Emotional States
A non-invasive method for the monitoring of heart activity can help to reduce the deaths caused by heart disorders such as stroke, arrhythmia and heart attack. The human voice can be considered as a biometric data that c…
Extending RNN-T-based speech recognition systems with emotion and language classification
Speech transcription, emotion recognition, and language identification are usually considered to be three different tasks. Each one requires a different model with a different architecture and training process. We propos…
Emotion ClassificationEmotion RecognitionLanguage Identificationspeech-recognition+2Best Practices for Noise-Based Augmentation to Improve the Performance of Deployable Speech-Based Emotion Recognition Systems
Speech emotion recognition is an important component of any human centered system. But speech characteristics produced and perceived by a person can be influenced by a multitude of reasons, both desirable such as emotion…
Adversarial AttackAutomatic Speech RecognitionData AugmentationEmotion Recognition+4Multi-stream Attention-based BLSTM with Feature Segmentation for Speech Emotion Recognition
This paper proposes a speech emotion recognition technique that considers the suprasegmental characteristics and temporal change of individual speech parameters. In recent years, speech emotion recognition using Bidir…
Data AugmentationEmotional Speech SynthesisEmotion RecognitionSpeech Emotion Recognition+1EEG based Emotion Recognition of Image Stimuli
Emotion is playing a great role in our daily lives. The necessity and importance of an automatic Emotion recognition system is getting increased. Traditional approaches of emotion recognition are based on facial images, …
EEGElectroencephalogram (EEG)Emotion RecognitionMarketing