paper-with-me

홈 › Papers

GLA-Grad++: An Improved Griffin-Lim Guided Diffusion Model for Speech Synthesis

2025-11-27 · Teysir Baoueb, Xiaoyu Bie, Mathieu Fontaine, Gaël Richard arxiv

Recent advances in diffusion models have positioned them as powerful generative frameworks for speech synthesis, demonstrating substantial improvements in audio quality and stability. Nevertheless, their effectiveness in vocoders conditioned on mel spectrograms remains constrained, particularly when the conditioning diverges from the training distribution. The recently proposed GLA-Grad model introduced a phase-aware extension to the WaveGrad vocoder that integrated the Griffin-Lim algorithm (GLA) into the reverse process to reduce inconsistencies between generated signals and conditioning mel spectrogram. In this paper, we further improve GLA-Grad through an innovative choice in how to apply the correction. Particularly, we compute the correction term only once, with a single application of GLA, to accelerate the generation process. Experimental results demonstrate that our method consistently outperforms the baseline models, particularly in out-of-domain scenarios.

📄 PDF Abstract BibTeX arXiv:2511.22293

Code (0)

등록된 구현이 없습니다.

Tasks

Speech Synthesis

Similar Papers 제목 키워드 기반

GLA-Grad: A Griffin-Lim Extended Waveform Generation Diffusion Model

2024-02-09 · Haocheng Liu, Teysir Baoueb, Mathieu Fontaine, Jonathan Le Roux 외

Diffusion models are receiving a growing interest for a variety of signal generation tasks such as speech or music synthesis. WaveGrad, for example, is a successful diffusion model that conditionally uses the mel spectro…

Beyond Griffin-Lim: Improved Iterative Phase Retrieval for Speech

2022-05-11 · Tal Peer, Simon Welker, Timo Gerkmann

Phase retrieval is a problem encountered not only in speech and audio processing, but in many other fields such as optics. Iterative algorithms based on non-convex set projections are effective and frequently used for re…

Retrieval

A Flexible Online Framework for Projection-Based STFT Phase Retrieval

2023-09-13 · Tal Peer, Simon Welker, Johannes Kolhoff, Timo Gerkmann

Several recent contributions in the field of iterative STFT phase retrieval have demonstrated that the performance of the classical Griffin-Lim method can be considerably improved upon. By using the same projection opera…

Retrieval

High-quality Speech Synthesis Using Super-resolution Mel-Spectrogram

2019-12-03 · Leyuan Sheng, Dong-Yan Huang, Evgeniy N. Pavlovskiy

In speech synthesis and speech enhancement systems, melspectrograms need to be precise in acoustic representations. However, the generated spectrograms are over-smooth, that could not produce high quality synthesized spe…

Image-to-Image TranslationSpeech EnhancementSpeech SynthesisSuper-Resolution+2

Guided-TTS: A Diffusion Model for Text-to-Speech via Classifier Guidance

2021-11-23 · Heeseung Kim, Sungwon Kim, Sungroh Yoon

We propose Guided-TTS, a high-quality text-to-speech (TTS) model that does not require any transcript of target speaker using classifier guidance. Guided-TTS combines an unconditional diffusion probabilistic model with a…

speech-recognitionSpeech RecognitionSpeech Synthesistext-to-speech+2