paper-with-me

홈 › Papers

Robust Multi-channel Speech Recognition using Frequency Aligned Network

2020-02-06 · Taejin Park, Kenichi Kumatani, Minhua Wu, Shiva Sundaram

Conventional speech enhancement technique such as beamforming has known benefits for far-field speech recognition. Our own work in frequency-domain multi-channel acoustic modeling has shown additional improvements by training a spatial filtering layer jointly within an acoustic model. In this paper, we further develop this idea and use frequency aligned network for robust multi-channel automatic speech recognition (ASR). Unlike an affine layer in the frequency domain, the proposed frequency aligned component prevents one frequency bin influencing other frequency bins. We show that this modification not only reduces the number of parameters in the model but also significantly and improves the ASR performance. We investigate effects of frequency aligned network through ASR experiments on the real-world far-field data where users are interacting with an ASR system in uncontrolled acoustic environments. We show that our multi-channel acoustic model with a frequency aligned network shows up to 18% relative reduction in word error rate.

📄 PDF Abstract BibTeX arXiv:2002.02520

Code (0)

등록된 구현이 없습니다.

Tasks

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Speech Enhancementspeech-recognitionSpeech Recognition

Similar Papers 제목 키워드 기반

A comprehensive study of speech separation: spectrogram vs waveform separation

2019-05-17 · Fahimeh Bahmaninezhad, Jian Wu, Rongzhi Gu, Shi-Xiong Zhang 외

Speech separation has been studied widely for single-channel close-talk microphone recordings over the past few years; developed solutions are mostly in frequency-domain. Recently, a raw audio waveform separation network…

speech-recognitionSpeech RecognitionSpeech Separation

Noise-Robust ASR for the third 'CHiME' Challenge Exploiting Time-Frequency Masking based Multi-Channel Speech Enhancement and Recurrent Neural Network

2015-09-24 · Zaihu Pang, Fengyun Zhu

In this paper, the Lingban entry to the third 'CHiME' speech separation and recognition challenge is presented. A time-frequency masking based speech enhancement front-end is proposed to suppress the environmental noise …

Language ModelingLanguage ModellingSpeech Enhancementspeech-recognition+2

3-D Feature and Acoustic Modeling for Far-Field Speech Recognition

2019-11-13 · Anurenjan Purushothaman, Anirudh Sreeram, Sriram Ganapathy

Automatic speech recognition in multi-channel reverberant conditions is a challenging task. The conventional way of suppressing the reverberation artifacts involves a beamforming based enhancement of the multi-channel sp…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognitionSpeech Recognition

Rank-1 Constrained Multichannel Wiener Filter for Speech Recognition in Noisy Environments

2017-07-01 · Ziteng Wang, Emmanuel Vincent, Romain Serizel, Yonghong Yan

Multichannel linear filters, such as the Multichannel Wiener Filter (MWF) and the Generalized Eigenvalue (GEV) beamformer are popular signal processing techniques which can improve speech recognition performance. In this…

speech-recognitionSpeech Recognition

Searching for Effective Preprocessing Method and CNN-based Architecture with Efficient Channel Attention on Speech Emotion Recognition

2024-09-06 · Byunggun Kim, Younghun Kwon

Speech emotion recognition (SER) classifies human emotions in speech with a computer model. Recently, performance in SER has steadily increased as deep learning techniques have adapted. However, unlike many domains that …

Emotion RecognitionSpeech Emotion Recognition