paper-with-me

홈 › Papers

Variance-Preserving-Based Interpolation Diffusion Models for Speech Enhancement

2023-06-14 · Zilu Guo, Jun Du, Chin-Hui Lee, Yu Gao, Wenbin Zhang

The goal of this study is to implement diffusion models for speech enhancement (SE). The first step is to emphasize the theoretical foundation of variance-preserving (VP)-based interpolation diffusion under continuous conditions. Subsequently, we present a more concise framework that encapsulates both the VP- and variance-exploding (VE)-based interpolation diffusion methods. We demonstrate that these two methods are special cases of the proposed framework. Additionally, we provide a practical example of VP-based interpolation diffusion for the SE task. To improve performance and ease model training, we analyze the common difficulties encountered in diffusion models and suggest amenable hyper-parameters. Finally, we evaluate our model against several methods using a public benchmark to showcase the effectiveness of our approach

📄 PDF Abstract BibTeX arXiv:2306.08527

Code (1)

zelokuo/VPIDM 공식 구현 pytorch

Tasks

Speech Enhancement

Methods 이 논문이 사용한 방법론

Diffusion Diffusion models generate samples by gradually removing noise from a signal, and their training objective can be expressed as a reweighted variational lower-bound…

Similar Papers 제목 키워드 기반

A Variance-Preserving Interpolation Approach for Diffusion Models with Applications to Single Channel Speech Enhancement and Recognition

2024-05-27 · Zilu Guo, Qing Wang, Jun Du, Jia Pan 외

In this paper, we propose a variance-preserving interpolation framework to improve diffusion models for single-channel speech enhancement (SE) and automatic speech recognition (ASR). This new variance-preserving interpol…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Speech Enhancementspeech-recognition+1

An Analysis of the Variance of Diffusion-based Speech Enhancement

2024-02-01 · Bunlong Lay, Timo Gerkmann

Diffusion models proved to be powerful models for generative speech enhancement. In recent SGMSE+ approaches, training involves a stochastic differential equation for the diffusion process, adding both Gaussian and envir…

Speech Enhancement

Normalize Everything: A Preconditioned Magnitude-Preserving Architecture for Diffusion-Based Speech Enhancement

2025-05-08 · Julius Richter, Danilo de Oliveira, Timo Gerkmann

This paper presents a new framework for diffusion-based speech enhancement. Our method employs a Schroedinger bridge to transform the noisy speech distribution into the clean speech distribution. To stabilize and improve…

Image GenerationSpeech Enhancement

Frame Interpolation with Consecutive Brownian Bridge Diffusion

2024-05-09 · Zonglin Lyu, Ming Li, Jianbo Jiao, Chen Chen

Recent work in Video Frame Interpolation (VFI) tries to formulate VFI as a diffusion-based conditional image generation problem, synthesizing the intermediate frame given a random noise and neighboring frames. Due to the…

Conditional Image GenerationImage GenerationVideo Frame Interpolation

ArtiFree: Detecting and Reducing Generative Artifacts in Diffusion-based Speech Enhancement

2025-09-23 · Bhawana Chhaglani, Yang Gao, Julius Richter, Xilin Li 외 arxiv

Diffusion-based speech enhancement (SE) achieves natural-sounding speech and strong generalization, yet suffers from key limitations like generative artifacts and high inference latency. In this work, we systematically s…

Speech Enhancement