paper-with-me

홈 › Papers

High Fidelity Speech Enhancement with Band-split RNN

2022-12-01 · Jianwei Yu, Yi Luo, Hangting Chen, Rongzhi Gu, Chao Weng

Despite the rapid progress in speech enhancement (SE) research, enhancing the quality of desired speech in environments with strong noise and interfering speakers remains challenging. In this paper, we extend the application of the recently proposed band-split RNN (BSRNN) model to full-band SE and personalized SE (PSE) tasks. To mitigate the effects of unstable high-frequency components in full-band speech, we perform bi-directional and uni-directional band-level modeling to low-frequency and high-frequency subbands, respectively. For PSE task, we incorporate a speaker enrollment module into BSRNN to utilize target speaker information. Moreover, we utilize a MetricGAN discriminator (MGD) and a multi-resolution spectrogram discriminator (MRSD) to improve perceptual quality metrics. Experimental results show that our system outperforms various top-ranking SE systems, achieves state-of-the-art (SOTA) results on the DNS-2020 test set and ranks among the top 3 in the DNS-2023 challenge.

📄 PDF Abstract BibTeX arXiv:2212.00406

Code (1)

sungwon23/bsrnn pytorch

Tasks

Speech EnhancementVocal Bursts Intensity Prediction

Methods 이 논문이 사용한 방법론

Test 설명 없음

Similar Papers 제목 키워드 기반

Speech Bandwidth Expansion Via High Fidelity Generative Adversarial Networks

2024-07-26 · Mahmoud Salhab, Haidar Harmanani

Speech bandwidth expansion is crucial for expanding the frequency range of low-bandwidth speech signals, thereby improving audio quality, clarity and perceptibility in digital applications. Its applications span telephon…

Generative Adversarial NetworkSpeech Enhancementspeech-recognitionSpeech Recognition+4

DiTSE: High-Fidelity Generative Speech Enhancement via Latent Diffusion Transformers

2025-04-13 · Heitor R. Guimarães, Jiaqi Su, Rithesh Kumar, Tiago H. Falk 외

Real-world speech recordings suffer from degradations such as background noise and reverberation. Speech enhancement aims to mitigate these issues by generating clean high-fidelity signals. While recent generative approa…

HallucinationSpeech Enhancement

Parallel Time-Band Mixing with Learned Observation-Adding for Robust ASR Front-Ends

2026-08-31 · Xingyu Shen, Runze Wang, Wei-Ping Zhu, Benoit Champagne arxiv

Speech enhancement is often used as a front-end for robust ASR, yet recurrent temporal and cross-band modules introduce sequential dependencies that reduce parallel efficiency. In this paper, we present a sequence-parall…

Speech Enhancement

Personalized speech enhancement combining band-split RNN and speaker attentive module

2023-02-20 · Xiaohuai Le, Li Chen, Chao He, Yiqing Guo 외

Target speaker information can be utilized in speech enhancement (SE) models to more effectively extract the desired speech. Previous works introduce the speaker embedding into speech enhancement models by means of conca…

Speech Enhancement

A two-step backward compatible fullband speech enhancement system

2022-01-26 · Xu Zhang, LianWu Chen, Xiguang Zheng, Xinlei Ren 외

Speech enhancement methods based on deep learning have surpassed traditional methods. While many of these new approaches are operating on the wideband (16kHz) sample rate, a new fullband (48kHz) speech enhancement system…

Speech EnhancementVocal Bursts Valence Prediction