paper-with-me

Papers

Multi-reference Tacotron by Intercross Training for Style Disentangling,Transfer and Control in Speech Synthesis

2019-04-04 · Yanyao Bian, Changbin Chen, Yongguo Kang, Zhenglin Pan

Speech style control and transfer techniques aim to enrich the diversity and expressiveness of synthesized speech. Existing approaches model all speech styles into one representation, lacking the ability to control a specific speech feature independently. To address this issue, we introduce a novel multi-reference structure to Tacotron and propose intercross training approach, which together ensure that each sub-encoder of the multi-reference encoder independently disentangles and controls a specific style. Experimental results show that our model is able to control and transfer desired speech styles individually.

📄 PDF Abstract BibTeX arXiv:1904.02373

Code (0)

등록된 구현이 없습니다.

Tasks

DiversitySpeech Synthesis

Methods 이 논문이 사용한 방법론

Griffin-Lim Algorithm The Griffin-Lim Algorithm (GLA) is a phase reconstruction method based on the redundancy of the short-time Fourier transform. It promotes the consistency of a spectrogram by…
Sigmoid Activation 설명 없음
Highway Layer 설명 없음
Residual Connection 설명 없음
Convolution A convolution is a type of matrix operation, consisting of a kernel, a small matrix of weights, that slides over input data performing element-wise multiplication with the…
Batch Normalization 설명 없음
Max Pooling Max Pooling is a pooling operation that calculates the maximum value for patches of a feature map, and uses it to create a downsampled (pooled) feature map. It is usually…
Residual GRU A Residual GRU is a gated recurrent unit (GRU) that incorporates the idea of residual connections from…

Similar Papers 제목 키워드 기반

Predicting Expressive Speaking Style From Text In End-To-End Speech Synthesis

2018-08-04 · Daisy Stanton, Yuxuan Wang, RJ Skerry-Ryan

Global Style Tokens (GSTs) are a recently-proposed method to learn latent disentangled representations of high-dimensional data. GSTs can be used within Tacotron, a state-of-the-art end-to-end text-to-speech synthesis sy…

Speech Synthesistext-to-speechText to SpeechText-To-Speech Synthesis

End-to-End Emotional Speech Synthesis Using Style Tokens and Semi-Supervised Training

2019-06-26

This paper proposes an end-to-end emotional speech synthesis (ESS) method which adopts global style tokens (GSTs) for semi-supervised training. This model is built based on the GST-Tacotron framework. The style tokens ar…

Emotional Speech SynthesisEmotion RecognitionSpeech Synthesis

Expressive Text-to-Speech using Style Tag

2021-04-01 · Minchan Kim, Sung Jun Cheon, Byoung Jin Choi, Jong Jin Kim 외

As recent text-to-speech (TTS) systems have been rapidly improved in speech quality and generation speed, many researchers now focus on a more challenging issue: expressive TTS. To control speaking styles, existing expre…

Language ModelingLanguage ModellingTAGtext-to-speech+1

Whispered and Lombard Neural Speech Synthesis

2021-01-13 · Qiong Hu, Tobias Bleisch, Petko Petkov, Tuomo Raitio 외

It is desirable for a text-to-speech system to take into account the environment where synthetic speech is presented, and provide appropriate context-dependent output to the user. In this paper, we present and compare va…

Speaker VerificationSpeech Synthesistext-to-speechText to Speech

Towards End-to-End Prosody Transfer for Expressive Speech Synthesis with Tacotron

2018-03-24 · ICML 2018 7 · RJ Skerry-Ryan, Eric Battenberg, Ying Xiao, Yuxuan Wang 외

We present an extension to the Tacotron speech synthesis architecture that learns a latent embedding space of prosody, derived from a reference acoustic representation containing the desired prosody. We show that conditi…

Expressive Speech SynthesisSpeech Synthesis