paper-with-me

홈 › Papers

Using VAEs and Normalizing Flows for One-shot Text-To-Speech Synthesis of Expressive Speech

2019-11-28 · Vatsal Aggarwal, Marius Cotescu, Nishant Prateek, Jaime Lorenzo-Trueba, Roberto Barra-Chicote

We propose a Text-to-Speech method to create an unseen expressive style using one utterance of expressive speech of around one second. Specifically, we enhance the disentanglement capabilities of a state-of-the-art sequence-to-sequence based system with a Variational AutoEncoder (VAE) and a Householder Flow. The proposed system provides a 22% KL-divergence reduction while jointly improving perceptual metrics over state-of-the-art. At synthesis time we use one example of expressive style as a reference input to the encoder for generating any text in the desired style. Perceptual MUSHRA evaluations show that we can create a voice with a 9% relative naturalness improvement over standard Neural Text-to-Speech, while also improving the perceived emotional intensity (59 compared to the 55 of neutral speech).

📄 PDF Abstract BibTeX arXiv:1911.12760

Code (0)

등록된 구현이 없습니다.

Tasks

DisentanglementExpressive Speech SynthesisSpeech Synthesistext-to-speechText to SpeechText-To-Speech Synthesis

Methods 이 논문이 사용한 방법론

Solana Customer Service Number +1-833-534-1729 설명 없음

Similar Papers 제목 키워드 기반

SurVAE Flows: Surjections to Bridge the Gap between VAEs and Flows

2020-07-06 · NeurIPS 2020 12 · Didrik Nielsen, Priyank Jaini, Emiel Hoogeboom, Ole Winther 외

Normalizing flows and variational autoencoders are powerful generative models that can represent complicated density functions. However, they both impose constraints on the models: Normalizing flows use bijective transfo…

Creating New Voices using Normalizing Flows

2023-12-22 · Piotr Bilinski, Thomas Merritt, Abdelhamid Ezzerg, Kamil Pokora 외

Creating realistic and natural-sounding synthetic speech remains a big challenge for voice identities unseen during training. As there is growing interest in synthesizing voices of new speakers, here we investigate the a…

Speech Synthesistext-to-speechText to SpeechVoice Conversion

CaloFlow: Fast and Accurate Generation of Calorimeter Showers with Normalizing Flows

2021-06-09 · Claudius Krause, David Shih

We introduce CaloFlow, a fast detector simulation framework based on normalizing flows. For the first time, we demonstrate that normalizing flows can reproduce many-channel calorimeter showers with extremely high fidelit…

Model Selection

Latent Variable Modelling with Hyperbolic Normalizing Flows

2020-02-15 · ICML 2020 1 · Avishek Joey Bose, Ariella Smofsky, Renjie Liao, Prakash Panangaden 외

The choice of approximate posterior distributions plays a central role in stochastic variational inference (SVI). One effective solution is the use of normalizing flows \cut{defined on Euclidean spaces} to construct flex…

Density EstimationVariational Inference

Generative Modeling for Low Dimensional Speech Attributes with Neural Spline Flows

2022-03-03 · Kevin J. Shih, Rafael Valle, Rohan Badlani, João Felipe Santos 외

Despite recent advances in generative modeling for text-to-speech synthesis, these models do not yet have the same fine-grained adjustability of pitch-conditioned deterministic models such as FastPitch and FastSpeech2. P…

Speech Synthesistext-to-speechText to SpeechText-To-Speech Synthesis