Emotional End-to-End Neural Speech Synthesizer
In this paper, we introduce an emotional speech synthesizer based on the recent end-to-end neural model, named Tacotron. Despite its benefits, we found that the original Tacotron suffers from the exposure bias problem and irregularity of the attention alignment. Later, we address the problem by utilization of context vector and residual connection at recurrent neural networks (RNNs). Our experiments showed that the model could successfully train and generate speech for given emotion labels.
Code (1)
Methods 이 논문이 사용한 방법론
Similar Papers 제목 키워드 기반
Multi-speaker Emotional Text-to-speech Synthesizer
We present a methodology to train our multi-speaker emotional text-to-speech synthesizer that can express speech for 10 speakers' 7 different emotions. All silences from audio samples are removed prior to learning. This …
Alltext-to-speechText to SpeechTransformer-Based Speech Synthesizer Attribution in an Open Set Scenario
Speech synthesis methods can create realistic-sounding speech, which may be used for fraud, spoofing, and misinformation campaigns. Forensic methods that detect synthesized speech are important for protection against suc…
AttributeMisinformationMulti-class ClassificationSpeech SynthesisDialogueAgents: A Hybrid Agent-Based Speech Synthesis Framework for Multi-Party Dialogue
Speech synthesis is crucial for human-computer interaction, enabling natural and intuitive communication. However, existing datasets involve high construction costs due to manual annotation and suffer from limited charac…
DiversitySpeech SynthesisSyntAct: A Synthesized Database of Basic Emotions
Speech emotion recognition is in the focus of research since several decades and has many applications. One problem is sparse data for supervised learning. One way to tackle this problem is the synthesis of data with emo…
Emotion RecognitionSpeech Emotion RecognitionSpeech SynthesisDiffusion Synthesizer for Efficient Multilingual Speech to Speech Translation
We introduce DiffuseST, a low-latency, direct speech-to-speech translation system capable of preserving the input speaker's voice zero-shot while translating from multiple source languages into English. We experiment wit…
Speech-to-Speech TranslationTranslation