paper-with-me

Papers

MetricGAN-U: Unsupervised speech enhancement/ dereverberation based only on noisy/ reverberated speech

2021-10-12 · Szu-Wei Fu, Cheng Yu, Kuo-Hsuan Hung, Mirco Ravanelli, Yu Tsao

Most of the deep learning-based speech enhancement models are learned in a supervised manner, which implies that pairs of noisy and clean speech are required during training. Consequently, several noisy speeches recorded in daily life cannot be used to train the model. Although certain unsupervised learning frameworks have also been proposed to solve the pair constraint, they still require clean speech or noise for training. Therefore, in this paper, we propose MetricGAN-U, which stands for MetricGAN-unsupervised, to further release the constraint from conventional unsupervised learning. In MetricGAN-U, only noisy speech is required to train the model by optimizing non-intrusive speech quality metrics. The experimental results verified that MetricGAN-U outperforms baselines in both objective and subjective metrics.

📄 PDF Abstract BibTeX arXiv:2110.05866

Code (2)

speechbrain/speechbrain/tree/develop/recipes/Voicebank 공식 구현 pytorch
xiaoyubie1994/dvae_se pytorch

Tasks

Speech Enhancement

Similar Papers 제목 키워드 기반

MetricGAN+: An Improved Version of MetricGAN for Speech Enhancement

2021-04-08 · Szu-Wei Fu, Cheng Yu, Tsun-An Hsieh, Peter Plantinga 외

The discrepancy between the cost function used for training a speech enhancement model and human auditory perception usually makes the quality of enhanced speech unsatisfactory. Objective evaluation metrics which conside…

Speech Enhancement

MetricGAN-OKD: Multi-Metric Optimization of MetricGAN via Online Knowledge Distillation for Speech Enhancement

2023-07-24 · ICML 2023 7 · WooSeok Shin, Byung Hoon Lee, Jin Sob Kim, Hyun Joon Park 외

In speech enhancement, MetricGAN-based approaches reduce the discrepancy between the $L_p$ loss and evaluation metrics by utilizing a non-differentiable evaluation metric as the objective function. However, optimizing mu…

Knowledge DistillationSpeech Enhancement

MetricGAN: Generative Adversarial Networks based Black-box Metric Scores Optimization for Speech Enhancement

2019-05-13 · Szu-Wei Fu, Chien-Feng Liao, Yu Tsao, Shou-De Lin

Adversarial loss in a conditional generative adversarial network (GAN) is not designed to directly optimize evaluation metrics of a target task, and thus, may not always guide the generator in a GAN to generate data with…

Generative Adversarial NetworkSpeech Enhancement

Boosting Objective Scores of a Speech Enhancement Model by MetricGAN Post-processing

2020-06-18 · Szu-Wei Fu, Chien-Feng Liao, Tsun-An Hsieh, Kuo-Hsuan Hung 외

The Transformer architecture has demonstrated a superior ability compared to recurrent neural networks in many different natural language processing applications. Therefore, our study applies a modified Transformer in a …

Speech Enhancement

BSS-CFFMA: Cross-Domain Feature Fusion and Multi-Attention Speech Enhancement Network based on Self-Supervised Embedding

2024-08-13 · Alimjan Mattursun, Liejun Wang, Yinfeng Yu

Speech self-supervised learning (SSL) represents has achieved state-of-the-art (SOTA) performance in multiple downstream tasks. However, its application in speech enhancement (SE) tasks remains immature, offering opportu…

DenoisingSelf-Supervised LearningSpeech Enhancement