paper-with-me

Papers

Training speech emotion classifier without categorical annotations

2022-10-14 · Meysam Shamsi, Marie Tahon

There are two paradigms of emotion representation, categorical labeling and dimensional description in continuous space. Therefore, the emotion recognition task can be treated as a classification or regression. The main aim of this study is to investigate the relation between these two representations and propose a classification pipeline that uses only dimensional annotation. The proposed approach contains a regressor model which is trained to predict a vector of continuous values in dimensional representation for given speech audio. The output of this model can be interpreted as an emotional category using a mapping algorithm. We investigated the performances of a combination of three feature extractors, three neural network architectures, and three mapping algorithms on two different corpora. Our study shows the advantages and limitations of the classification via regression approach.

📄 PDF Abstract BibTeX arXiv:2210.07642

Code (0)

등록된 구현이 없습니다.

Tasks

ClassificationEmotion Recognitionregression

Similar Papers 제목 키워드 기반

MSP-Podcast SER Challenge 2024: L'antenne du Ventoux Multimodal Self-Supervised Learning for Speech Emotion Recognition

2024-07-08 · Jarod Duret, Mickael Rouvier, Yannick Estève

In this work, we detail our submission to the 2024 edition of the MSP-Podcast Speech Emotion Recognition (SER) Challenge. This challenge is divided into two distinct tasks: Categorical Emotion Recognition and Emotional A…

AttributeEmotion RecognitionSelf-Supervised LearningSpeech Emotion Recognition

Learning Arousal-Valence Representation from Categorical Emotion Labels of Speech

2023-11-24 · Enting Zhou, You Zhang, Zhiyao Duan

Dimensional representations of speech emotions such as the arousal-valence (AV) representation provide a continuous and fine-grained description and control than their categorical counterparts. They have wide application…

Dimensionality ReductionEmotion ClassificationregressionSpeech Synthesis+3

MEDUSA: A Multimodal Deep Fusion Multi-Stage Training Framework for Speech Emotion Recognition in Naturalistic Conditions

2025-06-11 · Georgios Chatzichristodoulou, Despoina Kosmopoulou, Antonios Kritikos, Anastasia Poulopoulou 외

SER is a challenging task due to the subjective nature of human emotions and their uneven representation under naturalistic conditions. We propose MEDUSA, a multimodal framework with a four-stage training pipeline, which…

Emotion RecognitionSpeech Emotion Recognition

Detecting Emotion Primitives from Speech and their use in discerning Categorical Emotions

2020-01-31 · Vasudha Kowtha, Vikramjit Mitra, Chris Bartels, Erik Marchi 외

Emotion plays an essential role in human-to-human communication, enabling us to convey feelings such as happiness, frustration, and sincerity. While modern speech technologies rely heavily on speech recognition and natur…

Natural Language Understandingspeech-recognitionSpeech Recognition

Prompt-Unseen-Emotion: Zero-shot Expressive Speech Synthesis with Prompt-LLM Contextual Knowledge for Mixed Emotions

2025-06-03 · Xiaoxue Gao, Huayun Zhang, Nancy F. Chen

Existing expressive text-to-speech (TTS) systems primarily model a limited set of categorical emotions, whereas human conversations extend far beyond these predefined emotions, making it essential to explore more diverse…

Expressive Speech SynthesisPrompt LearningSpeech Synthesistext-to-speech+1