paper-with-me

홈 › Papers

Learning latent representations for style control and transfer in end-to-end speech synthesis

2018-12-11 · Ya-Jie Zhang, Shifeng Pan, Lei He, Zhen-Hua Ling

In this paper, we introduce the Variational Autoencoder (VAE) to an end-to-end speech synthesis model, to learn the latent representation of speaking styles in an unsupervised manner. The style representation learned through VAE shows good properties such as disentangling, scaling, and combination, which makes it easy for style control. Style transfer can be achieved in this framework by first inferring style representation through the recognition network of VAE, then feeding it into TTS network to guide the style in synthesizing speech. To avoid Kullback-Leibler (KL) divergence collapse in training, several techniques are adopted. Finally, the proposed model shows good performance of style control and outperforms Global Style Token (GST) model in ABX preference tests on style transfer.

📄 PDF Abstract BibTeX arXiv:1812.04342

Code (2)

jinhan/tacotron2-vae pytorch
yanggeng1995/vae_tacotron tf

Tasks

Speech SynthesisStyle Transfer

Methods 이 논문이 사용한 방법론

Solana Customer Service Number +1-833-534-1729 설명 없음
USD Coin Customer Service Number +1-833-534-1729 설명 없음

Similar Papers 제목 키워드 기반

Cross-speaker Emotion Transfer by Manipulating Speech Style Latents

2023-03-15 · Suhee Jo, Younggun Lee, Yookyung Shin, Yeongtae Hwang 외

In recent years, emotional text-to-speech has shown considerable progress. However, it requires a large amount of labeled data, which is not easily accessible. Even if it is possible to acquire an emotional speech datase…

text-to-speechText to Speech

Counterfactuals to Control Latent Disentangled Text Representations for Style Transfer

2021-08-01 · ACL 2021 5 · Sharmila Reddy Nangi, Niyati Chhaya, Sopan Khosla, Nikhil Kaushik 외

Disentanglement of latent representations into content and style spaces has been a commonly employed method for unsupervised text style transfer. These techniques aim to learn the disentangled representations and tweak t…

AttributecounterfactualDisentanglementSentence+3

TCSinger: Zero-Shot Singing Voice Synthesis with Style Transfer and Multi-Level Style Control

2024-09-24 · Yu Zhang, Ziyue Jiang, RuiQi Li, Changhao Pan 외

Zero-shot singing voice synthesis (SVS) with style transfer and style control aims to generate high-quality singing voices with unseen timbres and styles (including singing method, emotion, rhythm, technique, and pronunc…

ClusteringLanguage ModellingQuantizationRhythm+2

Expressive Neural Voice Cloning

2021-01-30 · Paarth Neekhara, Shehzeen Hussain, Shlomo Dubnov, Farinaz Koushanfar 외

Voice cloning is the task of learning to synthesize the voice of an unseen speaker from a few samples. While current voice cloning methods achieve promising results in Text-to-Speech (TTS) synthesis for a new voice, thes…

Speech SynthesisStyle Transfertext-to-speechText to Speech+1

An Empirical Study on Learning Latent Representations for Emotional Speech Synthesis

2026-06-12 · Vinh Dang Quang, Huy Ngo Quang arxiv

For the last couple of years, the field of speech synthesis has improved dramatically thanks to deep learning. There are more and more deep learning-based TTS systems developed to make it possible to produce voices with …

Speech Synthesis