paper-with-me

홈 › Papers

An Investigation of End-to-End Models for Robust Speech Recognition

2021-02-11 · Archiki Prasad, Preethi Jyothi, Rajbabu Velmurugan

End-to-end models for robust automatic speech recognition (ASR) have not been sufficiently well-explored in prior work. With end-to-end models, one could choose to preprocess the input speech using speech enhancement techniques and train the model using enhanced speech. Another alternative is to pass the noisy speech as input and modify the model architecture to adapt to noisy speech. A systematic comparison of these two approaches for end-to-end robust ASR has not been attempted before. We address this gap and present a detailed comparison of speech enhancement-based techniques and three different model-based adaptation techniques covering data augmentation, multi-task learning, and adversarial learning for robust ASR. While adversarial learning is the best-performing technique on certain noise types, it comes at the cost of degrading clean speech WER. On other relatively stationary noise types, a new speech enhancement technique outperformed all the model-based adaptation techniques. This suggests that knowledge of the underlying noise type can meaningfully inform the choice of adaptation technique.

📄 PDF Abstract BibTeX arXiv:2102.06237

Code (1)

archiki/Robust-E2E-ASR 공식 구현 pytorch

Tasks

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Data AugmentationMulti-Task LearningRobust Speech RecognitionSpeech Enhancementspeech-recognitionSpeech Recognition

Similar Papers 제목 키워드 기반

A Comprehensive Study of the Current State-of-the-Art in Nepali Automatic Speech Recognition Systems

2024-02-05 · Rupak Raj Ghimire, Bal Krishna Bal, Prakash Poudyal

In this paper, we examine the research conducted in the field of Nepali Automatic Speech Recognition (ASR). The primary objective of this survey is to conduct a comprehensive review of the works on Nepali Automatic Speec…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognitionSpeech Recognition

Investigations on End-to-End Audiovisual Fusion

2018-04-30 · Michael Wand, Ngoc Thang Vu, Juergen Schmidhuber

Audiovisual speech recognition (AVSR) is a method to alleviate the adverse effect of noise in the acoustic signal. Leveraging recent developments in deep neural network-based speech recognition, we present an AVSR neural…

speech-recognitionSpeech Recognition

探究端對端混合模型架構於華語語音辨識 (An Investigation of Hybrid CTC-Attention Modeling in Mandarin Speech Recognition)

2019-06-01 · IJCLCLP 2019 6 · Hsiu-jui Chang, Wei-Cheng Chao, Tien-Hong Lo, Berlin Chen
speech-recognitionSpeech Recognition

Qualitative investigation of the display of speech recognition results for communication with deaf people

2015-09-01 · WS 2015 9 · Agn{\`e}s Piquard-Kipffer, Odile Mella, Mir, J{\'e}r{\'e}my a 외
Language Modellingspeech-recognitionSpeech Recognition

End-to-End Multi-Speaker Speech Recognition using Speaker Embeddings and Transfer Learning

2019-08-13 · Pavel Denisov, Ngoc Thang Vu

This paper presents our latest investigation on end-to-end automatic speech recognition (ASR) for overlapped speech. We propose to train an end-to-end system conditioned on speaker embeddings and further improved by tran…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognitionSpeech Recognition+1