paper-with-me

홈 › Papers

Emotional End-to-End Neural Speech Synthesizer

2017-11-15 · Young-Gun Lee, Azam Rabiee, Soo-Young Lee

In this paper, we introduce an emotional speech synthesizer based on the recent end-to-end neural model, named Tacotron. Despite its benefits, we found that the original Tacotron suffers from the exposure bias problem and irregularity of the attention alignment. Later, we address the problem by utilization of context vector and residual connection at recurrent neural networks (RNNs). Our experiments showed that the model could successfully train and generate speech for given emotion labels.

📄 PDF Abstract BibTeX arXiv:1711.05447

Code (1)

AzamRabiee/Emotional-TTS 공식 구현

Methods 이 논문이 사용한 방법론

Griffin-Lim Algorithm The Griffin-Lim Algorithm (GLA) is a phase reconstruction method based on the redundancy of the short-time Fourier transform. It promotes the consistency of a spectrogram by…
Sigmoid Activation 설명 없음
Highway Layer 설명 없음
Convolution A convolution is a type of matrix operation, consisting of a kernel, a small matrix of weights, that slides over input data performing element-wise multiplication with the…
Batch Normalization 설명 없음
Max Pooling Max Pooling is a pooling operation that calculates the maximum value for patches of a feature map, and uses it to create a downsampled (pooled) feature map. It is usually…
Residual GRU A Residual GRU is a gated recurrent unit (GRU) that incorporates the idea of residual connections from…
BiGRU A Bidirectional GRU, or BiGRU, is a sequence processing model that consists of two GRUs. one taking the input in a forward…

Similar Papers 제목 키워드 기반

Multi-speaker Emotional Text-to-speech Synthesizer

2021-12-07 · Sungjae Cho, Soo-Young Lee

We present a methodology to train our multi-speaker emotional text-to-speech synthesizer that can express speech for 10 speakers' 7 different emotions. All silences from audio samples are removed prior to learning. This …

Alltext-to-speechText to Speech

Transformer-Based Speech Synthesizer Attribution in an Open Set Scenario

2022-10-14 · Emily R. Bartusiak, Edward J. Delp

Speech synthesis methods can create realistic-sounding speech, which may be used for fraud, spoofing, and misinformation campaigns. Forensic methods that detect synthesized speech are important for protection against suc…

AttributeMisinformationMulti-class ClassificationSpeech Synthesis

DialogueAgents: A Hybrid Agent-Based Speech Synthesis Framework for Multi-Party Dialogue

2025-04-20 · Xiang Li, Duyi Pan, Hongru Xiao, Jiale Han 외

Speech synthesis is crucial for human-computer interaction, enabling natural and intuitive communication. However, existing datasets involve high construction costs due to manual annotation and suffer from limited charac…

DiversitySpeech Synthesis

SyntAct: A Synthesized Database of Basic Emotions

2022-06-01 · DCLRL (LREC) 2022 6 · Felix Burkhardt, Florian Eyben, Björn Schuller

Speech emotion recognition is in the focus of research since several decades and has many applications. One problem is sparse data for supervised learning. One way to tackle this problem is the synthesis of data with emo…

Emotion RecognitionSpeech Emotion RecognitionSpeech Synthesis

Diffusion Synthesizer for Efficient Multilingual Speech to Speech Translation

2024-06-14 · Nameer Hirschkind, Xiao Yu, Mahesh Kumar Nandwana, Joseph Liu 외

We introduce DiffuseST, a low-latency, direct speech-to-speech translation system capable of preserving the input speaker's voice zero-shot while translating from multiple source languages into English. We experiment wit…

Speech-to-Speech TranslationTranslation