paper-with-me

Papers

WaveGrad 2: Iterative Refinement for Text-to-Speech Synthesis

2021-06-17 · Nanxin Chen, Yu Zhang, Heiga Zen, Ron J. Weiss, Mohammad Norouzi, Najim Dehak, William Chan

This paper introduces WaveGrad 2, a non-autoregressive generative model for text-to-speech synthesis. WaveGrad 2 is trained to estimate the gradient of the log conditional density of the waveform given a phoneme sequence. The model takes an input phoneme sequence, and through an iterative refinement process, generates an audio waveform. This contrasts to the original WaveGrad vocoder which conditions on mel-spectrogram features, generated by a separate model. The iterative refinement process starts from Gaussian noise, and through a series of refinement steps (e.g., 50 steps), progressively recovers the audio sequence. WaveGrad 2 offers a natural way to trade-off between inference speed and sample quality, through adjusting the number of refinement steps. Experiments show that the model can generate high fidelity audio, approaching the performance of a state-of-the-art neural TTS system. We also report various ablation studies over different model configurations. Audio samples are available at https://wavegrad.github.io/v2.

📄 PDF Abstract BibTeX arXiv:2106.09660

Code (3)

keonlee9420/WaveGrad2 pytorch
maum-ai/wavegrad2 pytorch
mindslab-ai/wavegrad2 pytorch

Tasks

Speech Synthesistext-to-speechText to SpeechText-To-Speech Synthesis

Methods 이 논문이 사용한 방법론

HuMan(Expedia)||How do I get a human at Expedia? How do I get a human at Expedia? How Do I Get a Human at Expedia? – Call ☎️ +1-(888) 829 (0881) or +1-805-330-4056 or +1-805-330-4056 for Real-Time Help & Exclusive…
Residual Connection 설명 없음
1x1 Convolution A 1 x 1 Convolution is a convolution with some special properties in that it can be used for dimensionality reduction,…
WaveGrad UBlock 설명 없음
FiLM Module 설명 없음
Convolution A convolution is a type of matrix operation, consisting of a kernel, a small matrix of weights, that slides over input data performing element-wise multiplication with the…
WaveGrad DBlock WaveGrad DBlocks are used to downsample the temporal dimension of noisy waveform in WaveGrad. They are similar to UBlocks except…
WaveGrad WaveGrad is a conditional model for waveform generation through estimating gradients of the data density. This model is built on the prior work on score matching and diffusion…

Similar Papers 제목 키워드 기반

WaveGrad: Estimating Gradients for Waveform Generation

2020-09-02 · ICLR 2021 1 · Nanxin Chen, Yu Zhang, Heiga Zen, Ron J. Weiss 외

This paper introduces WaveGrad, a conditional model for waveform generation which estimates gradients of the data density. The model is built on prior work on score matching and diffusion probabilistic models. It starts …

Speech SynthesisText-To-Speech Synthesis

Multi-interaction TTS toward professional recording reproduction

2025-07-01 · Hiroki Kanagawa, Kenichi Fujita, Aya Watanabe, Yusuke Ijima arxiv

Voice directors often iteratively refine voice actors' performances by providing feedback to achieve the desired outcome. While this iterative feedback-based refinement process is important in actual recordings, it has b…

Text-To-Speech Synthesis

GLA-Grad++: An Improved Griffin-Lim Guided Diffusion Model for Speech Synthesis

2025-11-27 · Teysir Baoueb, Xiaoyu Bie, Mathieu Fontaine, Gaël Richard arxiv

Recent advances in diffusion models have positioned them as powerful generative frameworks for speech synthesis, demonstrating substantial improvements in audio quality and stability. Nevertheless, their effectiveness in…

Speech Synthesis

SpecGrad: Diffusion Probabilistic Model based Neural Vocoder with Adaptive Noise Spectral Shaping

2022-03-31 · Yuma Koizumi, Heiga Zen, Kohei Yatabe, Nanxin Chen 외

Neural vocoder using denoising diffusion probabilistic model (DDPM) has been improved by adaptation of the diffusion noise distribution to given acoustic features. In this study, we propose SpecGrad that adapts the diffu…

DenoisingSpeech Enhancement

A Non-autoregressive Model for Joint STT and TTS

2025-01-15 · Vishal Sunder, Brian Kingsbury, George Saon, Samuel Thomas 외

In this paper, we take a step towards jointly modeling automatic speech recognition (STT) and speech synthesis (TTS) in a fully non-autoregressive way. We develop a novel multimodal framework capable of handling the spee…

Automatic Speech Recognitionspeech-recognitionSpeech RecognitionSpeech Synthesis