paper-with-me

Papers

SpecGrad: Diffusion Probabilistic Model based Neural Vocoder with Adaptive Noise Spectral Shaping

2022-03-31 · Yuma Koizumi, Heiga Zen, Kohei Yatabe, Nanxin Chen, Michiel Bacchiani

Neural vocoder using denoising diffusion probabilistic model (DDPM) has been improved by adaptation of the diffusion noise distribution to given acoustic features. In this study, we propose SpecGrad that adapts the diffusion noise so that its time-varying spectral envelope becomes close to the conditioning log-mel spectrogram. This adaptation by time-varying filtering improves the sound quality especially in the high-frequency bands. It is processed in the time-frequency domain to keep the computational cost almost the same as the conventional DDPM-based neural vocoders. Experimental results showed that SpecGrad generates higher-fidelity speech waveform than conventional DDPM-based neural vocoders in both analysis-synthesis and speech enhancement scenarios. Audio demos are available at wavegrad.github.io/specgrad/.

📄 PDF Abstract BibTeX arXiv:2203.16749

Code (0)

등록된 구현이 없습니다.

Tasks

DenoisingSpeech Enhancement

Methods 이 논문이 사용한 방법론

Diffusion Diffusion models generate samples by gradually removing noise from a signal, and their training objective can be expressed as a reweighted variational lower-bound…

Similar Papers 제목 키워드 기반

PeriodGrad: Towards Pitch-Controllable Neural Vocoder Based on a Diffusion Probabilistic Model

2024-02-22 · Yukiya Hono, Kei Hashimoto, Yoshihiko Nankaku, Keiichi Tokuda

This paper presents a neural vocoder based on a denoising diffusion probabilistic model (DDPM) incorporating explicit periodic signals as auxiliary conditioning signals. Recently, DDPM-based neural vocoders have gained p…

DenoisingPitch controlSinging Voice Synthesis

PriorGrad: Improving Conditional Denoising Diffusion Models with Data-Dependent Adaptive Prior

2021-06-11 · ICLR 2022 4 · Sang-gil Lee, Heeseung Kim, Chaehun Shin, Xu Tan 외

Denoising diffusion probabilistic models have been recently proposed to generate high-quality samples by estimating the gradient of the data density. The framework defines the prior noise as a standard Gaussian distribut…

Audio GenerationDenoisingSpeech SynthesisText-To-Speech Synthesis

WaveFit: An Iterative and Non-autoregressive Neural Vocoder based on Fixed-Point Iteration

2022-10-03 · Yuma Koizumi, Kohei Yatabe, Heiga Zen, Michiel Bacchiani

Denoising diffusion probabilistic models (DDPMs) and generative adversarial networks (GANs) are popular generative models for neural vocoders. The DDPMs and GANs can be characterized by the iterative denoising framework …

Denoising

NU-Wave: A Diffusion Probabilistic Model for Neural Audio Upsampling

2021-04-06 · Junhyeok Lee, Seungu Han

In this work, we introduce NU-Wave, the first neural audio upsampling model to produce waveforms of sampling rate 48kHz from coarse 16kHz or 24kHz inputs, while prior works could generate only up to 16kHz. NU-Wave is the…

Audio Super-ResolutionSuper-Resolution

InferGrad: Improving Diffusion Models for Vocoder by Considering Inference in Training

2022-02-08 · Zehua Chen, Xu Tan, Ke Wang, Shifeng Pan 외

Denoising diffusion probabilistic models (diffusion models for short) require a large number of iterations in inference to achieve the generation quality that matches or surpasses the state-of-the-art generative models, …

Denoising