paper-with-me

홈 › Papers

GLA-Grad: A Griffin-Lim Extended Waveform Generation Diffusion Model

2024-02-09 · Haocheng Liu, Teysir Baoueb, Mathieu Fontaine, Jonathan Le Roux, Gael Richard

Diffusion models are receiving a growing interest for a variety of signal generation tasks such as speech or music synthesis. WaveGrad, for example, is a successful diffusion model that conditionally uses the mel spectrogram to guide a diffusion process for the generation of high-fidelity audio. However, such models face important challenges concerning the noise diffusion process for training and inference, and they have difficulty generating high-quality speech for speakers that were not seen during training. With the aim of minimizing the conditioning error and increasing the efficiency of the noise diffusion process, we propose in this paper a new scheme called GLA-Grad, which consists in introducing a phase recovery algorithm such as the Griffin-Lim algorithm (GLA) at each step of the regular diffusion process. Furthermore, it can be directly applied to an already-trained waveform generation model, without additional training or fine-tuning. We show that our algorithm outperforms state-of-the-art diffusion models for speech generation, especially when generating speech for a previously unseen target speaker.

📄 PDF Abstract BibTeX arXiv:2402.15516

Code (0)

등록된 구현이 없습니다.

Methods 이 논문이 사용한 방법론

Convolution A convolution is a type of matrix operation, consisting of a kernel, a small matrix of weights, that slides over input data performing element-wise multiplication with the…
HuMan(Expedia)||How do I get a human at Expedia? How do I get a human at Expedia? How Do I Get a Human at Expedia? – Call ☎️ +1-(888) 829 (0881) or +1-805-330-4056 or +1-805-330-4056 for Real-Time Help & Exclusive…
Diffusion Diffusion models generate samples by gradually removing noise from a signal, and their training objective can be expressed as a reweighted variational lower-bound…
1x1 Convolution A 1 x 1 Convolution is a convolution with some special properties in that it can be used for dimensionality reduction,…
Residual Connection 설명 없음
Griffin-Lim Algorithm The Griffin-Lim Algorithm (GLA) is a phase reconstruction method based on the redundancy of the short-time Fourier transform. It promotes the consistency of a spectrogram by…
FiLM Module 설명 없음
WaveGrad DBlock WaveGrad DBlocks are used to downsample the temporal dimension of noisy waveform in WaveGrad. They are similar to UBlocks except…

Similar Papers 제목 키워드 기반

GLA-Grad++: An Improved Griffin-Lim Guided Diffusion Model for Speech Synthesis

2025-11-27 · Teysir Baoueb, Xiaoyu Bie, Mathieu Fontaine, Gaël Richard arxiv

Recent advances in diffusion models have positioned them as powerful generative frameworks for speech synthesis, demonstrating substantial improvements in audio quality and stability. Nevertheless, their effectiveness in…

Speech Synthesis

Fast Spectrogram Inversion using Multi-head Convolutional Neural Networks

2018-08-20 · Sercan O. Arik, Heewoo Jun, Gregory Diamos

We propose the multi-head convolutional neural network (MCNN) architecture for waveform synthesis from spectrograms. Nonlinear interpolation in MCNN is employed with transposed convolution layers in parallel heads. MCNN …

speech-recognitionSpeech RecognitionSpeech Synthesis

Predicting Different Acoustic Features from EEG and towards direct synthesis of Audio Waveform from EEG

2020-05-29 · Gautam Krishna, Co Tran, Mason Carnahan, Ahmed Tewfik

In [1,2] authors provided preliminary results for synthesizing speech from electroencephalography (EEG) features where they first predict acoustic features from EEG features and then the speech is reconstructed from the …

EEGElectroencephalogram (EEG)

Prompt-prompted Adaptive Structured Pruning for Efficient LLM Generation

2024-04-01 · Harry Dong, Beidi Chen, Yuejie Chi

With the development of transformer-based large language models (LLMs), they have been applied to many fields due to their remarkable utility, but this comes at a considerable computational cost at deployment. Fortunatel…

Mixture-of-Experts

UnDiff: Unsupervised Voice Restoration with Unconditional Diffusion Model

2023-06-01 · Anastasiia Iashchenko, Pavel Andreev, Ivan Shchekotov, Nicholas Babaev 외

This paper introduces UnDiff, a diffusion probabilistic model capable of solving various speech inverse tasks. Being once trained for speech waveform generation in an unconditional manner, it can be adapted to different …

Bandwidth Extensionmodel