paper-with-me

Papers

Exploring Multilingual Unseen Speaker Emotion Recognition: Leveraging Co-Attention Cues in Multitask Learning

2024-06-13 · Arnav Goel, Medha Hira, Anubha Gupta

Advent of modern deep learning techniques has given rise to advancements in the field of Speech Emotion Recognition (SER). However, most systems prevalent in the field fail to generalize to speakers not seen during training. This study focuses on handling challenges of multilingual SER, specifically on unseen speakers. We introduce CAMuLeNet, a novel architecture leveraging co-attention based fusion and multitask learning to address this problem. Additionally, we benchmark pretrained encoders of Whisper, HuBERT, Wav2Vec2.0, and WavLM using 10-fold leave-speaker-out cross-validation on five existing multilingual benchmark datasets: IEMOCAP, RAVDESS, CREMA-D, EmoDB and CaFE and, release a novel dataset for SER on the Hindi language (BhavVani). CAMuLeNet shows an average improvement of approximately 8% over all benchmarks on unseen speakers determined by our cross-validation strategy.

📄 PDF Abstract BibTeX arXiv:2406.08931

Code (1)

arnav10goel/camulenet 공식 구현

Tasks

Emotion RecognitionSpeech Emotion Recognition

Similar Papers 제목 키워드 기반

Large Language Models Meet Contrastive Learning: Zero-Shot Emotion Recognition Across Languages

2025-03-25 · Heqing Zou, Fengmao Lv, Desheng Zheng, Eng Siong Chng 외

Multilingual speech emotion recognition aims to estimate a speaker's emotional state using a contactless method across different languages. However, variability in voice characteristics and linguistic diversity poses sig…

Contrastive LearningDiversityEmotion RecognitionSpeech Emotion Recognition

Effect of Attention and Self-Supervised Speech Embeddings on Non-Semantic Speech Tasks

2023-08-28 · Payal Mohapatra, Akash Pandey, Yueyuan Sui, Qi Zhu

Human emotion understanding is pivotal in making conversational technology mainstream. We view speech emotion understanding as a perception task which is a more realistic setting. With varying contexts (languages, demogr…

Speech Recognition

Speaker-invariant Affective Representation Learning via Adversarial Training

2019-11-04 · Haoqi Li, Ming Tu, Jing Huang, Shrikanth Narayanan 외

Representation learning for speech emotion recognition is challenging due to labeled data sparsity issue and lack of gold standard references. In addition, there is much variability from input speech signals, human subje…

Emotion ClassificationEmotion RecognitionRepresentation LearningSpeech Emotion Recognition

Seen and Unseen emotional style transfer for voice conversion with a new emotional speech dataset

2020-10-28 · Kun Zhou, Berrak Sisman, Rui Liu, Haizhou Li

Emotional voice conversion aims to transform emotional prosody in speech while preserving the linguistic content and speaker identity. Prior studies show that it is possible to disentangle emotional prosody using an enco…

DecoderEmotion RecognitionGenerative Adversarial NetworkSpeech Emotion Recognition+2

EM2LDL: A Multilingual Speech Corpus for Mixed Emotion Recognition through Label Distribution Learning

2025-11-25 · Xingfeng Li, Xiaohan Shi, Junjie Li, Yongwei Li 외 arxiv

This study introduces EM2LDL, a novel multilingual speech corpus designed to advance mixed emotion recognition through label distribution learning. Addressing the limitations of predominantly monolingual and single-label…

Self-Supervised LearningEmotion Recognition