paper-with-me

홈 › Papers

Thunder : Unified Regression-Diffusion Speech Enhancement with a Single Reverse Step using Brownian Bridge

2024-06-10 · Thanapat Trachu, Chawan Piansaddhayanon, Ekapol Chuangsuwanich

Diffusion-based speech enhancement has shown promising results, but can suffer from a slower inference time. Initializing the diffusion process with the enhanced audio generated by a regression-based model can be used to reduce the computational steps required. However, these approaches often necessitate a regression model, further increasing the system's complexity. We propose Thunder, a unified regression-diffusion model that utilizes the Brownian bridge process which can allow the model to act in both modes. The regression mode can be accessed by setting the diffusion time step closed to 1. However, the standard score-based diffusion modeling does not perform well in this setup due to gradient instability. To mitigate this problem, we modify the diffusion model to predict the clean speech instead of the score function, achieving competitive performance with a more compact model size and fewer reverse steps.

📄 PDF Abstract BibTeX arXiv:2406.06139

Code (0)

등록된 구현이 없습니다.

Tasks

regressionSpeech Enhancement

Methods 이 논문이 사용한 방법론

Diffusion Diffusion models generate samples by gradually removing noise from a signal, and their training objective can be expressed as a reweighted variational lower-bound…

Similar Papers 제목 키워드 기반

uSee: Unified Speech Enhancement and Editing with Conditional Diffusion Models

2023-10-02 · Muqiao Yang, Chunlei Zhang, Yong Xu, Zhongweiyang Xu 외

Speech enhancement aims to improve the quality of speech signals in terms of quality and intelligibility, and speech editing refers to the process of editing the speech according to specific user needs. In this paper, we…

DenoisingSelf-Supervised LearningSpeech DenoisingSpeech Enhancement

ProSE: Diffusion Priors for Speech Enhancement

2025-03-09 · Sonal Kumar, Sreyan Ghosh, Utkarsh Tyagi, Anton Jeran Ratnarajah 외

Speech enhancement (SE) is the foundational task of enhancing the clarity and quality of speech in the presence of non-stationary additive noise. While deterministic deep learning models have been commonly employed for S…

DenoisingregressionSpeech Enhancement

Supervising 3D Talking Head Avatars with Analysis-by-Audio-Synthesis

2025-04-18 · Radek Daněček, Carolin Schmitt, Senya Polikovsky, Michael J. Black

In order to be widely applicable, speech-driven 3D head avatars must articulate their lips in accordance with speech, while also conveying the appropriate emotions with dynamically changing facial expressions. The key pr…

Audio Synthesis

Diff-SV: A Unified Hierarchical Framework for Noise-Robust Speaker Verification Using Score-Based Diffusion Probabilistic Models

2023-09-14 · Ju-ho Kim, Jungwoo Heo, Hyun-seo Shin, Chan-yeong Lim 외

Background noise considerably reduces the accuracy and reliability of speaker verification (SV) systems. These challenges can be addressed using a speech enhancement system as a front-end module. Recently, diffusion prob…

Speaker VerificationSpeech Enhancement

Spoken Speech Enhancement using EEG

2019-09-13 · Gautam Krishna, Co Tran, Yan Han, Mason Carnahan 외

In this paper we demonstrate spoken speech enhancement using electroencephalography (EEG) signals using a generative adversarial network (GAN) based model, gated recurrent unit (GRU) regression based model, temporal conv…

EEGElectroencephalogram (EEG)Generative Adversarial Networkregression+1