paper-with-me

홈 › Papers

Unifying the Discrete and Continuous Emotion labels for Speech Emotion Recognition

2022-10-29 · Roshan Sharma, Hira Dhamyal, Bhiksha Raj, Rita Singh

Traditionally, in paralinguistic analysis for emotion detection from speech, emotions have been identified with discrete or dimensional (continuous-valued) labels. Accordingly, models that have been proposed for emotion detection use one or the other of these label types. However, psychologists like Russell and Plutchik have proposed theories and models that unite these views, maintaining that these representations have shared and complementary information. This paper is an attempt to validate these viewpoints computationally. To this end, we propose a model to jointly predict continuous and discrete emotional attributes and show how the relationship between these can be utilized to improve the robustness and performance of emotion recognition tasks. Our approach comprises multi-task and hierarchical multi-task learning frameworks that jointly model the relationships between continuous-valued and discrete emotion labels. Experimental results on two widely used datasets (IEMOCAP and MSPPodcast) for speech-based emotion recognition show that our model results in statistically significant improvements in performance over strong baselines with non-unified approaches. We also demonstrate that using one type of label (discrete or continuous-valued) for training improves recognition performance in tasks that use the other type of label. Experimental results and reasoning for this approach (called the mismatched training approach) are also presented.

📄 PDF Abstract BibTeX arXiv:2210.16642

Code (0)

등록된 구현이 없습니다.

Tasks

Emotion RecognitionMulti-Task LearningSpeech Emotion Recognition

Similar Papers 제목 키워드 기반

UDDETTS: Unifying Discrete and Dimensional Emotions for Controllable Emotional Text-to-Speech

2025-05-15 · Jiaxuan Liu, ZhenHua Ling

Recent neural codec language models have made great progress in the field of text-to-speech (TTS), but controllable emotional TTS still faces many challenges. Traditional methods rely on predefined discrete emotion label…

Emotional Speech SynthesisLanguage ModelingLanguage ModellingSpeech Synthesis+2

Empirical Interpretation of the Relationship Between Speech Acoustic Context and Emotion Recognition

2023-06-30 · Anna Ollerenshaw, Md Asif Jalal, Rosanna Milner, Thomas Hain

Speech emotion recognition (SER) is vital for obtaining emotional intelligence and understanding the contextual meaning of speech. Variations of consonant-vowel (CV) phonemic boundaries can enrich acoustic context with l…

Emotional IntelligenceEmotion RecognitionSpeech Emotion Recognition

EmoSteer-TTS: Fine-Grained and Training-Free Emotion-Controllable Text-to-Speech via Activation Steering

2025-08-05 · Tianxin Xie, Shan Yang, Chenxing Li, Dong Yu 외 arxiv

Text-to-speech (TTS) has shown great progress in recent years. However, most existing TTS systems offer only coarse and rigid emotion control, typically via discrete emotion labels or a carefully crafted and detailed emo…

Continuous Control

Facial Expression Editing with Continuous Emotion Labels

2020-06-22 · Alexandra Lindt, Pablo Barros, Henrique Siqueira, Stefan Wermter

Recently deep generative models have achieved impressive results in the field of automated facial expression editing. However, the approaches presented so far presume a discrete representation of human emotions and are t…

EmoSphere-SER: Enhancing Speech Emotion Recognition Through Spherical Representation with Auxiliary Classification

2025-05-26 · Deok-Hyeon Cho, Hyung-Seok Oh, Seung-bin Kim, Seong-Whan Lee

Speech emotion recognition predicts a speaker's emotional state from speech signals using discrete labels or continuous dimensions such as arousal, valence, and dominance (VAD). We propose EmoSphere-SER, a joint model th…

Emotion RecognitionregressionSpeech Emotion Recognition