paper-with-me

Papers

End-to-End Far-Field Speech Recognition with Unified Dereverberation and Beamforming

2020-10-27

Despite successful applications of end-to-end approaches in multi-channel speech recognition, the performance still degrades severely when the speech is corrupted by reverberation. In this paper, we integrate the dereverberation module into the end-to-end multi-channel speech recognition system and explore two different frontend architectures. First, a multi-source mask-based weighted prediction error (WPE) module is incorporated in the frontend for dereverberation. Second, another novel frontend architecture is proposed, which extends the weighted power minimization distortionless response (WPD) convolutional beamformer to perform simultaneous separation and dereverberation. We derive a new formulation from the original WPD, which can handle multi-source input, and replace eigenvalue decomposition with the matrix inverse operation to make the back-propagation algorithm more stable. The above two architectures are optimized in a fully end-to-end manner, only using the speech recognition criterion. Experiments on both spatialized wsj1-2mix corpus and REVERB show that our proposed model outperformed the conventional methods in reverberant scenarios.

📄 PDF Abstract BibTeX arXiv:2005.10479

Code (0)

등록된 구현이 없습니다.

Tasks

speech-recognitionSpeech Recognition

Similar Papers 제목 키워드 기반

Run-Time Adaptation of Neural Beamforming for Robust Speech Dereverberation and Denoising

2024-10-30 · Yoto Fujita, Aditya Arie Nugraha, Diego Di Carlo, Yoshiaki Bando 외

This paper describes speech enhancement for realtime automatic speech recognition (ASR) in real environments. A standard approach to this task is to use neural beamforming that can work efficiently in an online manner. I…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)DenoisingSpeech Dereverberation+3

An Investigation of End-to-End Multichannel Speech Recognition for Reverberant and Mismatch Conditions

2019-04-19 · Aswin Shanmugam Subramanian, Xiaofei Wang, Shinji Watanabe, Toru Taniguchi 외

Sequence-to-sequence (S2S) modeling is becoming a popular paradigm for automatic speech recognition (ASR) because of its ability to jointly optimize all the conventional ASR components in an end-to-end (E2E) fashion. Thi…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)DenoisingNoisy Speech Recognition+3

End-to-End Dereverberation, Beamforming, and Speech Recognition with Improved Numerical Stability and Advanced Frontend

2021-02-23 · Wangyou Zhang, Christoph Boeddeker, Shinji Watanabe, Tomohiro Nakatani 외

Recently, the end-to-end approach has been successfully applied to multi-speaker speech separation and recognition in both single-channel and multichannel conditions. However, severe performance degradation is still obse…

Action DetectionActivity DetectionSpeech Dereverberationspeech-recognition+2

Blind Speech Separation and Dereverberation using Neural Beamforming

2021-03-24 · Lukas Pfeifenberger, Franz Pernkopf

In this paper, we present the Blind Speech Separation and Dereverberation (BSSD) network, which performs simultaneous speaker separation, dereverberation and speaker identification in a single neural network. Speaker sep…

Speaker IdentificationSpeaker SeparationSpeech SeparationTriplet

Dereverberation of Autoregressive Envelopes for Far-field Speech Recognition

2021-08-12 · Anurenjan Purushothaman, Anirudh Sreeram, Rohit Kumar, Sriram Ganapathy

The task of speech recognition in far-field environments is adversely affected by the reverberant artifacts that elicit as the temporal smearing of the sub-band envelopes. In this paper, we develop a neural model for spe…

Speech Dereverberationspeech-recognitionSpeech Recognition