paper-with-me

Papers

Contrastive Unsupervised Learning for Speech Emotion Recognition

2021-02-12 · Mao Li, Bo Yang, Joshua Levy, Andreas Stolcke, Viktor Rozgic, Spyros Matsoukas, Constantinos Papayiannis, Daniel Bone, Chao Wang

Speech emotion recognition (SER) is a key technology to enable more natural human-machine communication. However, SER has long suffered from a lack of public large-scale labeled datasets. To circumvent this problem, we investigate how unsupervised representation learning on unlabeled datasets can benefit SER. We show that the contrastive predictive coding (CPC) method can learn salient representations from unlabeled datasets, which improves emotion recognition performance. In our experiments, this method achieved state-of-the-art concordance correlation coefficient (CCC) performance for all emotion primitives (activation, valence, and dominance) on IEMOCAP. Additionally, on the MSP- Podcast dataset, our method obtained considerable performance improvements compared to baselines.

📄 PDF Abstract BibTeX arXiv:2102.06357

Code (0)

등록된 구현이 없습니다.

Tasks

Emotion RecognitionRepresentation LearningSpeech Emotion Recognition

Methods 이 논문이 사용한 방법론

InfoNCE 설명 없음
Contrastive Predictive Coding Contrastive Predictive Coding (CPC) learns self-supervised representations by predicting the future in latent space by using powerful autoregressive models. The model uses a…

Similar Papers 제목 키워드 기반

Non-Contrastive Self-Supervised Learning of Utterance-Level Speech Representations

2022-08-10 · Jaejin Cho, Raghavendra Pappagari, Piotr Żelasko, Laureano Moro-Velazquez 외

Considering the abundance of unlabeled speech data and the high labeling costs, unsupervised learning methods can be essential for better system development. One of the most successful methods is contrastive self-supervi…

Emotion RecognitionSelf-Supervised LearningSpeaker Verification

CochCeps-Augment: A Novel Self-Supervised Contrastive Learning Using Cochlear Cepstrum-based Masking for Speech Emotion Recognition

2024-02-10 · Ioannis Ziogas, Hessa Alfalahi, Ahsan H. Khandoker, Leontios J. Hadjileontiadis

Self-supervised learning (SSL) for automated speech recognition in terms of its emotional content, can be heavily degraded by the presence noise, affecting the efficiency of modeling the intricate temporal and spectral i…

Contrastive LearningEmotion RecognitionImage AugmentationSelf-Supervised Learning+4

Emo-DNA: Emotion Decoupling and Alignment Learning for Cross-Corpus Speech Emotion Recognition

2023-08-04 · Jiaxin Ye, Yujie Wei, Xin-Cheng Wen, Chenglong Ma 외

Cross-corpus speech emotion recognition (SER) seeks to generalize the ability of inferring speech emotion from a well-labeled corpus to an unlabeled one, which is a rather challenging task due to the significant discrepa…

Cross-corpusDomain AdaptationEmotion RecognitionSpeech Emotion Recognition+1

Large Language Models Meet Contrastive Learning: Zero-Shot Emotion Recognition Across Languages

2025-03-25 · Heqing Zou, Fengmao Lv, Desheng Zheng, Eng Siong Chng 외

Multilingual speech emotion recognition aims to estimate a speaker's emotional state using a contactless method across different languages. However, variability in voice characteristics and linguistic diversity poses sig…

Contrastive LearningDiversityEmotion RecognitionSpeech Emotion Recognition

Learning More with Less: Self-Supervised Approaches for Low-Resource Speech Emotion Recognition

2025-06-01 · Ziwei Gong, Pengyuan Shi, Kaan Donbekci, Lin Ai 외

Speech Emotion Recognition (SER) has seen significant progress with deep learning, yet remains challenging for Low-Resource Languages (LRLs) due to the scarcity of annotated data. In this work, we explore unsupervised le…

Contrastive LearningEmotion RecognitionSpeech Emotion Recognition