Emotion-Regularized Conditional Variational Autoencoder for Emotional Response Generation
This paper presents an emotion-regularized conditional variational autoencoder (Emo-CVAE) model for generating emotional conversation responses. In conventional CVAE-based emotional response generation, emotion labels are simply used as additional conditions in prior, posterior and decoder networks. Considering that emotion styles are naturally entangled with semantic contents in the language space, the Emo-CVAE model utilizes emotion labels to regularize the CVAE latent space by introducing an extra emotion prediction network. In the training stage, the estimated latent variables are required to predict the emotion labels and token sequences of the input responses simultaneously. Experimental results show that our Emo-CVAE model can learn a more informative and structured latent space than a conventional CVAE model and output responses with better content and emotion performance than baseline CVAE and sequence-to-sequence (Seq2Seq) models.
Code (0)
등록된 구현이 없습니다.
Tasks
DecoderResponse GenerationMethods 이 논문이 사용한 방법론
Similar Papers 제목 키워드 기반
EEG2Vec: Learning Affective EEG Representations via Variational Autoencoders
There is a growing need for sparse representational formats of human affective states that can be utilized in scenarios with limited computational memory resources. We explore whether representing neural data, in respons…
Edge-computingEEGElectroencephalogram (EEG)Disentanglement of Emotional Style and Speaker Identity for Expressive Voice Conversion
Expressive voice conversion performs identity conversion for emotional speakers by jointly converting speaker identity and emotional style. Due to the hierarchical structure of speech emotion, it is challenging to disent…
DisentanglementVoice ConversionEmoGene: Audio-Driven Emotional 3D Talking-Head Generation
Audio-driven talking-head generation is a crucial and useful technology for virtual human interaction and film-making. While recent advances have focused on improving image fidelity and lip synchronization, generating ac…
NeRFTalking Head GenerationAdversarial Auto-encoders for Speech Based Emotion Recognition
Recently, generative adversarial networks and adversarial autoencoders have gained a lot of attention in machine learning community due to their exceptional performance in tasks such as digit classification and face reco…
Emotion RecognitionFace RecognitionData-driven emotional body language generation for social robotics
In social robotics, endowing humanoid robots with the ability to generate bodily expressions of affect can improve human-robot interaction and collaboration, since humans attribute, and perhaps subconsciously anticipate,…
AttributeText Generation