paper-with-me

홈 › Papers

Learning Representations of Emotional Speech with Deep Convolutional Generative Adversarial Networks

2017-04-22 · Jonathan Chang, Stefan Scherer

Automatically assessing emotional valence in human speech has historically been a difficult task for machine learning algorithms. The subtle changes in the voice of the speaker that are indicative of positive or negative emotional states are often "overshadowed" by voice characteristics relating to emotional intensity or emotional activation. In this work we explore a representation learning approach that automatically derives discriminative representations of emotional speech. In particular, we investigate two machine learning strategies to improve classifier performance: (1) utilization of unlabeled data using a deep convolutional generative adversarial network (DCGAN), and (2) multitask learning. Within our extensive experiments we leverage a multitask annotated emotional corpus as well as a large unlabeled meeting corpus (around 100 hours). Our speaker-independent classification experiments show that in particular the use of unlabeled data in our investigations improves performance of the classifiers and both fully supervised baseline approaches are outperformed considerably. We improve the classification of emotional valence on a discrete 5-point scale to 43.88% and on a 3-point scale to 49.80%, which is competitive to state-of-the-art performance.

📄 PDF Abstract BibTeX arXiv:1705.02394

Code (0)

등록된 구현이 없습니다.

Tasks

BIG-bench Machine LearningGeneral ClassificationGenerative Adversarial NetworkRepresentation Learning

Similar Papers 제목 키워드 기반

Seen and Unseen emotional style transfer for voice conversion with a new emotional speech dataset

2020-10-28 · Kun Zhou, Berrak Sisman, Rui Liu, Haizhou Li

Emotional voice conversion aims to transform emotional prosody in speech while preserving the linguistic content and speaker identity. Prior studies show that it is possible to disentangle emotional prosody using an enco…

DecoderEmotion RecognitionGenerative Adversarial NetworkSpeech Emotion Recognition+2

ET-GAN: Cross-Language Emotion Transfer Based on Cycle-Consistent Generative Adversarial Networks

2019-05-27 · Xiaoqi Jia, Jianwei Tai, Hang Zhou, Yakai Li 외

Despite the remarkable progress made in synthesizing emotional speech from text, it is still challenging to provide emotion information to existing speech segments. Previous methods mainly rely on parallel data, and few …

Domain AdaptationGenerative Adversarial NetworkSpeech SynthesisTransfer Learning

StarGANv2-VC: A Diverse, Unsupervised, Non-parallel Framework for Natural-Sounding Voice Conversion

2021-07-21 · Yinghao Aaron Li, Ali Zare, Nima Mesgarani

We present an unsupervised non-parallel many-to-many voice conversion (VC) method using a generative adversarial network (GAN) called StarGAN v2. Using a combination of adversarial source classifier loss and perceptual l…

Generative Adversarial Networktext-to-speechText to SpeechVoice Conversion

Learning Emotional Representations from Imbalanced Speech Data for Speech Emotion Recognition and Emotional Text-to-Speech

2023-06-09 · Shijun Wang, Jón Guðnason, Damian Borth

Effective speech emotional representations play a key role in Speech Emotion Recognition (SER) and Emotional Text-To-Speech (TTS) tasks. However, emotional speech samples are more difficult and expensive to acquire compa…

Emotion RecognitionSpeech Emotion Recognitiontext-to-speechText to Speech

VAW-GAN for Disentanglement and Recomposition of Emotional Elements in Speech

2020-11-03 · Kun Zhou, Berrak Sisman, Haizhou Li

Emotional voice conversion (EVC) aims to convert the emotion of speech from one state to another while preserving the linguistic content and speaker identity. In this paper, we study the disentanglement and recomposition…

DecoderDisentanglementGenerative Adversarial NetworkVoice Conversion