A Methodology for Controlling the Emotional Expressiveness in Synthetic Speech -- a Deep Learning approach
In this project, we aim to build a Text-to-Speech system able to produce speech with a controllable emotional expressiveness. We propose a methodology for solving this problem in three main steps. The first is the collection of emotional speech data. We discuss the various formats of existing datasets and their usability in speech generation. The second step is the development of a system to automatically annotate data with emotion/expressiveness features. We compare several techniques using transfer learning to extract such a representation through other tasks and propose a method to visualize and interpret the correlation between vocal and emotional features. The third step is the development of a deep learning-based system taking text and emotion/expressiveness as input and producing speech as output. We study the impact of fine tuning from a neutral TTS towards an emotional TTS in terms of intelligibility and perception of the emotion.
Code (0)
등록된 구현이 없습니다.
Tasks
text-to-speechText to SpeechTransfer LearningSimilar Papers 제목 키워드 기반
Cross-speaker Emotion Transfer by Manipulating Speech Style Latents
In recent years, emotional text-to-speech has shown considerable progress. However, it requires a large amount of labeled data, which is not easily accessible. Even if it is possible to acquire an emotional speech datase…
text-to-speechText to SpeechAn Empirical Study on Learning Latent Representations for Emotional Speech Synthesis
For the last couple of years, the field of speech synthesis has improved dramatically thanks to deep learning. There are more and more deep learning-based TTS systems developed to make it possible to produce voices with …
Speech SynthesisHuLA: Prosody-Aware Anti-Spoofing with Multi-Task Learning for Expressive and Emotional Synthetic Speech
Current anti-spoofing systems remain vulnerable to expressive and emotional synthetic speech, since they rarely leverage prosody as a discriminative cue. Prosody is central to human expressiveness and emotion, and humans…
Self-Supervised LearningMulti-Task LearningSpoof DetectionDurFlex-EVC: Duration-Flexible Emotional Voice Conversion Leveraging Discrete Representations without Text Alignment
Emotional voice conversion (EVC) involves modifying various acoustic characteristics, such as pitch and spectral envelope, to match a desired emotional state while preserving the speaker's identity. Existing EVC methods …
DisentanglementSelf-Supervised LearningText to SpeechVoice ConversionCan Emotion Fool Anti-spoofing?
Traditional anti-spoofing focuses on models and datasets built on synthetic speech with mostly neutral state, neglecting diverse emotional variations. As a result, their robustness against high-quality, emotionally expre…
Emotion RecognitionSpeech Emotion Recognitiontext-to-speechText to Speech