paper-with-me

홈 › Papers

Speaker Normalization for Self-supervised Speech Emotion Recognition

2022-02-02 · Itai Gat, Hagai Aronowitz, Weizhong Zhu, Edmilson Morais, Ron Hoory

Large speech emotion recognition datasets are hard to obtain, and small datasets may contain biases. Deep-net-based classifiers, in turn, are prone to exploit those biases and find shortcuts such as speaker characteristics. These shortcuts usually harm a model's ability to generalize. To address this challenge, we propose a gradient-based adversary learning framework that learns a speech emotion recognition task while normalizing speaker characteristics from the feature representation. We demonstrate the efficacy of our method on both speaker-independent and speaker-dependent settings and obtain new state-of-the-art results on the challenging IEMOCAP dataset.

📄 PDF Abstract BibTeX arXiv:2202.01252

Code (0)

등록된 구현이 없습니다.

Tasks

Emotion RecognitionSpeech Emotion Recognition

Similar Papers 제목 키워드 기반

Cross-speaker Emotion Transfer Based on Speaker Condition Layer Normalization and Semi-Supervised Training in Text-To-Speech

2021-10-08 · Pengfei Wu, Junjie Pan, Chenchang Xu, Junhui Zhang 외

In expressive speech synthesis, there are high requirements for emotion interpretation. However, it is time-consuming to acquire emotional audio corpus for arbitrary speakers due to their deduction ability. In response t…

Emotion InterpretationExpressive Speech SynthesisSpeech Synthesistext-to-speech+1

DiEmo-TTS: Disentangled Emotion Representations via Self-Supervised Distillation for Cross-Speaker Emotion Transfer in Text-to-Speech

2025-05-26 · Deok-Hyeon Cho, Hyung-Seok Oh, Seung-bin Kim, Seong-Whan Lee

Cross-speaker emotion transfer in speech synthesis relies on extracting speaker-independent emotion embeddings for accurate emotion modeling without retaining speaker traits. However, existing timbre compression methods …

AttributeEmotional Speech SynthesisSpeech Synthesistext-to-speech+1

In-the-wild Speech Emotion Conversion Using Disentangled Self-Supervised Representations and Neural Vocoder-based Resynthesis

2023-06-02 · Navin Raj Prabhu, Nale Lehmann-Willenbrock, Timo Gerkmann

Speech emotion conversion aims to convert the expressed emotion of a spoken utterance to a target emotion while preserving the lexical information and the speaker's identity. In this work, we specifically focus on in-the…

Resynthesis

Converting Anyone's Voice: End-to-End Expressive Voice Conversion with a Conditional Diffusion Model

2024-05-02 · Zongyang Du, Junchen Lu, Kun Zhou, Lakshmish Kaushik 외

Expressive voice conversion (VC) conducts speaker identity conversion for emotional speakers by jointly converting speaker identity and emotional style. Emotional style modeling for arbitrary speakers in expressive VC ha…

DenoisingEmotion RecognitionSpeaker VerificationSpeech Emotion Recognition+1

Extracting speaker and emotion information from self-supervised speech models via channel-wise correlations

2022-10-15 · Themos Stafylakis, Ladislav Mosner, Sofoklis Kakouros, Oldrich Plchot 외

Self-supervised learning of speech representations from large amounts of unlabeled data has enabled state-of-the-art results in several speech processing tasks. Aggregating these speech representations across time is typ…

DescriptiveSelf-Supervised Learning