Emotion Style Transfer with a Specified Intensity Using Deep Reinforcement Learning
Text style transfer is a widely explored task in natural language generation which aims to change the stylistic properties of the text while retaining its style-independent content. In this work, we propose the task of emotion style transfer with a specified intensity in an unsupervised setting. The aim is to rewrite a given sentence, in any emotion, to a target emotion while also controlling the intensity of the target emotion. Emotions are gradient in nature, some words/phrases represent higher emotional intensity, while others represent lower intensity. In this task, we want to control this gradient nature of the emotion in the output. Additionally, we explore the issues with the existing datasets and address them. A novel BART-based model is proposed that is trained for the task by direct rewards. Unlike existing work, we bootstrap the BART model by training it to generate paraphrases so that it can explore lexical and syntactic diversity required for the output. Extensive automatic and human evaluations show the efficacy of our model in solving the problem.
Code (0)
등록된 구현이 없습니다.
Tasks
Deep Reinforcement LearningDiversityreinforcement-learningReinforcement Learning (RL)SentenceStyle TransferText GenerationText Style TransferMethods 이 논문이 사용한 방법론
Similar Papers 제목 키워드 기반
EmoDiff: Intensity Controllable Emotional Text-to-Speech with Soft-Label Guidance
Although current neural text-to-speech (TTS) models are able to generate high-quality speech, intensity controllable emotional TTS is still a challenging task. Most existing methods need external optimizations for intens…
Denoisingtext-to-speechText to SpeechAffectEcho: Speaker Independent and Language-Agnostic Emotion and Affect Transfer for Speech Synthesis
Affect is an emotional characteristic encompassing valence, arousal, and intensity, and is a crucial attribute for enabling authentic conversations. While existing text-to-speech (TTS) and speech-to-speech systems rely o…
AttributeSpeech Synthesistext-to-speechText to SpeechEmoSphere++: Emotion-Controllable Zero-Shot Text-to-Speech via Emotion-Adaptive Spherical Vector
Emotional text-to-speech (TTS) technology has achieved significant progress in recent years; however, challenges remain owing to the inherent complexity of emotions and limitations of the available emotional speech datas…
DecoderEmotional Speech Synthesistext-to-speechText to SpeechQI-TTS: Questioning Intonation Control for Emotional Speech Synthesis
Recent expressive text to speech (TTS) models focus on synthesizing emotional speech, but some fine-grained styles such as intonation are neglected. In this paper, we propose QI-TTS which aims to better transfer and cont…
Emotional Speech SynthesisSentenceSpeech Synthesistext-to-speech+1Cross-speaker Emotion Transfer by Manipulating Speech Style Latents
In recent years, emotional text-to-speech has shown considerable progress. However, it requires a large amount of labeled data, which is not easily accessible. Even if it is possible to acquire an emotional speech datase…
text-to-speechText to Speech