paper-with-me

홈 › Papers

3-D Feature and Acoustic Modeling for Far-Field Speech Recognition

2019-11-13 · Anurenjan Purushothaman, Anirudh Sreeram, Sriram Ganapathy

Automatic speech recognition in multi-channel reverberant conditions is a challenging task. The conventional way of suppressing the reverberation artifacts involves a beamforming based enhancement of the multi-channel speech signal, which is used to extract spectrogram based features for a neural network acoustic model. In this paper, we propose to extract features directly from the multi-channel speech signal using a multi variate autoregressive (MAR) modeling approach, where the correlations among all the three dimensions of time, frequency and channel are exploited. The MAR features are fed to a convolutional neural network (CNN) architecture which performs the joint acoustic modeling on the three dimensions. The 3-D CNN architecture allows the combination of multi-channel features that optimize the speech recognition cost compared to the traditional beamforming models that focus on the enhancement task. Experiments are conducted on the CHiME-3 and REVERB Challenge dataset using multi-channel reverberant speech. In these experiments, the proposed 3-D feature and acoustic modeling approach provides significant improvements over an ASR system trained with beamformed audio (average relative improvements of 10 % and 9 % in word error rates for CHiME-3 and REVERB Challenge datasets respectively.

📄 PDF Abstract BibTeX arXiv:1911.05504

Code (0)

등록된 구현이 없습니다.

Tasks

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognitionSpeech Recognition

Similar Papers 제목 키워드 기반

Frequency Domain Multi-channel Acoustic Modeling for Distant Speech Recognition

2019-04-28

Conventional far-field automatic speech recognition (ASR) systems typically employ microphone array techniques for speech enhancement in order to improve robustness against noise or reverberation. However, such speech en…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Distant Speech RecognitionSpeech Enhancement+2

Improving Speech Recognition for the Elderly: A New Corpus of Elderly Japanese Speech and Investigation of Acoustic Modeling for Speech Recognition

2020-05-01 · LREC 2020 5 · Meiko Fukuda, Hiromitsu Nishizaki, Yurie Iribe, Ryota Nishimura 외

In an aging society like Japan, a highly accurate speech recognition system is needed for use in electronic devices for the elderly, but this level of accuracy cannot be obtained using conventional speech recognition sys…

speech-recognitionSpeech Recognition

Dereverberation of Autoregressive Envelopes for Far-field Speech Recognition

2021-08-12 · Anurenjan Purushothaman, Anirudh Sreeram, Rohit Kumar, Sriram Ganapathy

The task of speech recognition in far-field environments is adversely affected by the reverberant artifacts that elicit as the temporal smearing of the sub-band envelopes. In this paper, we develop a neural model for spe…

Speech Dereverberationspeech-recognitionSpeech Recognition

Estimating Phoneme Class Conditional Probabilities from Raw Speech Signal using Convolutional Neural Networks

2013-04-03 · Dimitri Palaz, Ronan Collobert, Mathew Magimai. -Doss

In hybrid hidden Markov model/artificial neural networks (HMM/ANN) automatic speech recognition (ASR) system, the phoneme class conditional probabilities are estimated by first extracting acoustic features from the speec…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Phoneme Recognitionspeech-recognition+1

Robust Multi-channel Speech Recognition using Frequency Aligned Network

2020-02-06 · Taejin Park, Kenichi Kumatani, Minhua Wu, Shiva Sundaram

Conventional speech enhancement technique such as beamforming has known benefits for far-field speech recognition. Our own work in frequency-domain multi-channel acoustic modeling has shown additional improvements by tra…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Speech Enhancementspeech-recognition+1