paper-with-me

홈 › Papers

Batch-normalized joint training for DNN-based distant speech recognition

2017-03-24 · Mirco Ravanelli, Philemon Brakel, Maurizio Omologo, Yoshua Bengio

Improving distant speech recognition is a crucial step towards flexible human-machine interfaces. Current technology, however, still exhibits a lack of robustness, especially when adverse acoustic conditions are met. Despite the significant progress made in the last years on both speech enhancement and speech recognition, one potential limitation of state-of-the-art technology lies in composing modules that are not well matched because they are not trained jointly. To address this concern, a promising approach consists in concatenating a speech enhancement and a speech recognition deep neural network and to jointly update their parameters as if they were within a single bigger network. Unfortunately, joint training can be difficult because the output distribution of the speech enhancement system may change substantially during the optimization procedure. The speech recognition module would have to deal with an input distribution that is non-stationary and unnormalized. To mitigate this issue, we propose a joint training approach based on a fully batch-normalized architecture. Experiments, conducted using different datasets, tasks and acoustic conditions, revealed that the proposed framework significantly overtakes other competitive solutions, especially in challenging environments.

📄 PDF Abstract BibTeX arXiv:1703.08471

Code (0)

등록된 구현이 없습니다.

Tasks

Distant Speech RecognitionSpeech Enhancementspeech-recognitionSpeech Recognition

Similar Papers 제목 키워드 기반

End-to-end Joint Punctuated and Normalized ASR with a Limited Amount of Punctuated Training Data

2023-11-29 · Can Cui, Imran Ahamad Sheikh, Mostafa Sadeghi, Emmanuel Vincent

Joint punctuated and normalized automatic speech recognition (ASR), that outputs transcripts with and without punctuation and casing, remains challenging due to the lack of paired speech and punctuated text data in most …

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)DecoderLanguage Modeling+3

A network of deep neural networks for distant speech recognition

2017-03-23 · Mirco Ravanelli, Philemon Brakel, Maurizio Omologo, Yoshua Bengio

Despite the remarkable progress recently made in distant speech recognition, state-of-the-art technology still suffers from a lack of robustness, especially when adverse acoustic conditions characterized by non-stationar…

Distant Speech RecognitionSpeech Enhancementspeech-recognitionSpeech Recognition

Neural Blind Source Separation and Diarization for Distant Speech Recognition

2024-06-12 · Yoshiaki Bando, Tomohiko Nakamura, Shinji Watanabe

This paper presents a neural method for distant speech recognition (DSR) that jointly separates and diarizes speech mixtures without supervision by isolated signals. A standard separation method for multi-talker DSR is a…

blind source separationDistant Speech Recognitionspeaker-diarizationSpeaker Diarization+2

Normalization Before Shaking Toward Learning Symmetrically Distributed Representation Without Margin in Speech Emotion Recognition

2018-08-02 · Che-Wei Huang, Shrikanth. S. Narayanan

Regularization is crucial to the success of many practical deep learning models, in particular in a more often than not scenario where there are only a few to a moderate number of accessible training samples. In addition…

Data AugmentationEmotion RecognitionGeneral ClassificationSpeech Emotion Recognition

A Study of Enhancement, Augmentation, and Autoencoder Methods for Domain Adaptation in Distant Speech Recognition

2018-06-13 · Hao Tang, Wei-Ning Hsu, Francois Grondin, James Glass

Speech recognizers trained on close-talking speech do not generalize to distant speech and the word error rate degradation can be as large as 40% absolute. Most studies focus on tackling distant speech recognition as a s…

Data AugmentationDistant Speech RecognitionDomain AdaptationSpeech Enhancement+2