paper-with-me

Papers

DiffPhase: Generative Diffusion-based STFT Phase Retrieval

2022-11-08 · Tal Peer, Simon Welker, Timo Gerkmann

Diffusion probabilistic models have been recently used in a variety of tasks, including speech enhancement and synthesis. As a generative approach, diffusion models have been shown to be especially suitable for imputation problems, where missing data is generated based on existing data. Phase retrieval is inherently an imputation problem, where phase information has to be generated based on the given magnitude. In this work we build upon previous work in the speech domain, adapting a speech enhancement diffusion model specifically for STFT phase retrieval. Evaluation using speech quality and intelligibility metrics shows the diffusion approach is well-suited to the phase retrieval task, with performance surpassing both classical and modern methods.

📄 PDF Abstract BibTeX arXiv:2211.04332

Code (0)

등록된 구현이 없습니다.

Tasks

ImputationRetrievalSpeech Enhancement

Methods 이 논문이 사용한 방법론

Diffusion Diffusion models generate samples by gradually removing noise from a signal, and their training objective can be expressed as a reweighted variational lower-bound…

Similar Papers 제목 키워드 기반

Real-Time Streaming Mel Vocoding with Generative Flow Matching

2025-09-18 · Simon Welker, Tal Peer, Timo Gerkmann arxiv

The task of Mel vocoding, i.e., the inversion of a Mel magnitude spectrogram to an audio waveform, is still a key component in many text-to-speech (TTS) systems today. Based on generative flow matching, our prior work on…

Time-Frequency Phase Retrieval for Audio -- The Effect of Transform Parameters

2021-06-09 · Andrés Marafioti, Nicki Holighaus, Piotr Majdak

In audio processing applications, phase retrieval (PR) is often performed from the magnitude of short-time Fourier transform (STFT) coefficients. Although PR performance has been observed to depend on the considered STFT…

Retrieval

Real-Time Streamable Generative Speech Restoration with Flow Matching

2025-12-22 · Simon Welker, Bunlong Lay, Maris Hillemann, Tal Peer 외 arxiv

Diffusion-based generative models have greatly impacted the speech processing field in recent years, exhibiting high speech naturalness and spawning a new research direction. Their application in real-time communication …

Bandwidth ExtensionSpeech Enhancement

DOLPH: Diffusion Models for Phase Retrieval

2022-11-01 · Shirin Shoushtari, Jiaming Liu, Ulugbek S. Kamilov

Phase retrieval refers to the problem of recovering an image from the magnitudes of its complex-valued linear measurements. Since the problem is ill-posed, the recovery requires prior knowledge on the unknown image. We p…

Retrieval

Unsupervised speech enhancement with diffusion-based generative models

2023-09-19 · Berné Nortier, Mostafa Sadeghi, Romain Serizel

Recently, conditional score-based diffusion models have gained significant attention in the field of supervised speech enhancement, yielding state-of-the-art performance. However, these methods may face challenges when g…

Speech Enhancement