paper-with-me

홈 › Papers

CNN-based MultiChannel End-to-End Speech Recognition for everyday home environments

2018-11-07 · Nelson Yalta, Shinji Watanabe, Takaaki Hori, Kazuhiro Nakadai, Tetsuya OGATA

Casual conversations involving multiple speakers and noises from surrounding devices are common in everyday environments, which degrades the performances of automatic speech recognition systems. These challenging characteristics of environments are the target of the CHiME-5 challenge. By employing a convolutional neural network (CNN)-based multichannel end-to-end speech recognition system, this study attempts to overcome the presents difficulties in everyday environments. The system comprises of an attention-based encoder-decoder neural network that directly generates a text as an output from a sound input. The multichannel CNN encoder, which uses residual connections and batch renormalization, is trained with augmented data, including white noise injection. The experimental results show that the word error rate is reduced by 8.5% and 0.6% absolute from a single channel end-to-end and the best baseline (LF-MMI TDNN) on the CHiME-5 corpus, respectively.

📄 PDF Abstract BibTeX arXiv:1811.02735

Code (0)

등록된 구현이 없습니다.

Tasks

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Decoderspeech-recognitionSpeech Recognition

Similar Papers 제목 키워드 기반

Rank-1 Constrained Multichannel Wiener Filter for Speech Recognition in Noisy Environments

2017-07-01 · Ziteng Wang, Emmanuel Vincent, Romain Serizel, Yonghong Yan

Multichannel linear filters, such as the Multichannel Wiener Filter (MWF) and the Generalized Eigenvalue (GEV) beamformer are popular signal processing techniques which can improve speech recognition performance. In this…

speech-recognitionSpeech Recognition

CHiME-6 Challenge:Tackling Multispeaker Speech Recognition for Unsegmented Recordings

2020-04-20 · Shinji Watanabe, Michael Mandel, Jon Barker, Emmanuel Vincent 외

Following the success of the 1st, 2nd, 3rd, 4th and 5th CHiME challenges we organize the 6th CHiME Speech Separation and Recognition Challenge (CHiME-6). The new challenge revisits the previous CHiME-5 challenge and furt…

speaker-diarizationSpeaker DiarizationSpeech Enhancementspeech-recognition+2

Unsupervised Speech Enhancement Based on Multichannel NMF-Informed Beamforming for Noise-Robust Automatic Speech Recognition

2019-03-22 · Kazuki Shimada, Yoshiaki Bando, Masato Mimura, Katsutoshi Itoyama 외

This paper describes multichannel speech enhancement for improving automatic speech recognition (ASR) in noisy environments. Recently, the minimum variance distortionless response (MVDR) beamforming has widely been used …

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Speech Enhancementspeech-recognition+1

Multichannel AV-wav2vec2: A Framework for Learning Multichannel Multi-Modal Speech Representation

2024-01-07 · Qiushi Zhu, Jie Zhang, Yu Gu, Yuchen Hu 외

Self-supervised speech pre-training methods have developed rapidly in recent years, which show to be very effective for many near-field single-channel speech tasks. However, far-field multichannel speech processing is su…

Audio-Visual Speech RecognitionAutomatic Speech RecognitionAutomatic Speech Recognition (ASR)Contrastive Learning+6

CirdoX: an on/off-line multisource speech and sound analysis software

2016-05-01 · LREC 2016 5 · Fr{\'e}d{\'e}ric Aman, Michel Vacher, Fran{\c{c}}ois Portet, William Duclot 외

Vocal User Interfaces in domestic environments recently gained interest in the speech processing community. This interest is due to the opportunity of using it in the framework of Ambient Assisted Living both for home au…

General Classification