paper-with-me

Papers

Unsupervised Speech Enhancement with speech recognition embedding and disentanglement losses

2021-11-16 · Viet Anh Trinh, Sebastian Braun

Speech enhancement has recently achieved great success with various deep learning methods. However, most conventional speech enhancement systems are trained with supervised methods that impose two significant challenges. First, a majority of training datasets for speech enhancement systems are synthetic. When mixing clean speech and noisy corpora to create the synthetic datasets, domain mismatches occur between synthetic and real-world recordings of noisy speech or audio. Second, there is a trade-off between increasing speech enhancement performance and degrading speech recognition (ASR) performance. Thus, we propose an unsupervised loss function to tackle those two problems. Our function is developed by extending the MixIT loss function with speech recognition embedding and disentanglement loss. Our results show that the proposed function effectively improves the speech enhancement performance compared to a baseline trained in a supervised way on the noisy VoxCeleb dataset. While fully unsupervised training is unable to exceed the corresponding baseline, with joint super- and unsupervised training, the system is able to achieve similar speech quality and better ASR performance than the best supervised baseline.

📄 PDF Abstract BibTeX arXiv:2111.08678

Code (0)

등록된 구현이 없습니다.

Tasks

DisentanglementSpeech Enhancementspeech-recognitionSpeech Recognition

Similar Papers 제목 키워드 기반

A Closer Look at Wav2Vec2 Embeddings for On-Device Single-Channel Speech Enhancement

2024-03-03 · Ravi Shankar, Ke Tan, Buye Xu, Anurag Kumar

Self-supervised learned models have been found to be very effective for certain speech tasks such as automatic speech recognition, speaker identification, keyword spotting and others. While the features are undeniably us…

Automatic Speech RecognitionKeyword SpottingKnowledge DistillationSpeaker Identification+3

Speaker Re-identification with Speaker Dependent Speech Enhancement

2020-05-15 · Yanpei Shi, Qiang Huang, Thomas Hain

While the use of deep neural networks has significantly boosted speaker recognition performance, it is still challenging to separate speakers in poor acoustic environments. Here speech enhancement methods have traditiona…

Speaker RecognitionSpeech Enhancement

Scaling Speech Enhancement in Unseen Environments with Noise Embeddings

2018-10-26 · Gil Keren, Jing Han, Björn Schuller

We address the problem of speech enhancement generalisation to unseen environments by performing two manipulations. First, we embed an additional recording from the environment alone, and use this embedding to alter acti…

Speech Enhancementspeech-recognitionSpeech Recognition

NeuralEcho: A Self-Attentive Recurrent Neural Network For Unified Acoustic Echo Suppression And Speech Enhancement

2022-05-20 · Meng Yu, Yong Xu, Chunlei Zhang, Shi-Xiong Zhang 외

Acoustic echo cancellation (AEC) plays an important role in the full-duplex speech communication as well as the front-end speech enhancement for recognition in the conditions when the loudspeaker plays back. In this pape…

Acoustic echo cancellationSpeech Enhancementspeech-recognitionSpeech Recognition

On monoaural speech enhancement for automatic recognition of real noisy speech using mixture invariant training

2022-05-03 · Jisi Zhang, Catalin Zorila, Rama Doddipatla, Jon Barker

In this paper, we explore an improved framework to train a monoaural neural enhancement model for robust speech recognition. The designed training framework extends the existing mixture invariant training criterion to ex…

Robust Speech RecognitionSpeech Enhancementspeech-recognitionSpeech Recognition