paper-with-me

홈 › Papers

Multichannel End-to-end Speech Recognition

2017-03-14 · ICML 2017 8 · Tsubasa Ochiai, Shinji Watanabe, Takaaki Hori, John R. Hershey

The field of speech recognition is in the midst of a paradigm shift: end-to-end neural networks are challenging the dominance of hidden Markov models as a core technology. Using an attention mechanism in a recurrent encoder-decoder architecture solves the dynamic time alignment problem, allowing joint end-to-end training of the acoustic and language modeling components. In this paper we extend the end-to-end framework to encompass microphone array signal processing for noise suppression and speech enhancement within the acoustic encoding network. This allows the beamforming components to be optimized jointly within the recognition architecture to improve the end-to-end speech recognition objective. Experiments on the noisy speech benchmarks (CHiME-4 and AMI) show that our multichannel end-to-end system outperformed the attention-based baseline with input from a conventional adaptive beamformer.

📄 PDF Abstract BibTeX arXiv:1703.04783

Code (0)

등록된 구현이 없습니다.

Tasks

DecoderLanguage ModelingLanguage ModellingSpeech Enhancementspeech-recognitionSpeech Recognition

Similar Papers 제목 키워드 기반

Multichannel AV-wav2vec2: A Framework for Learning Multichannel Multi-Modal Speech Representation

2024-01-07 · Qiushi Zhu, Jie Zhang, Yu Gu, Yuchen Hu 외

Self-supervised speech pre-training methods have developed rapidly in recent years, which show to be very effective for many near-field single-channel speech tasks. However, far-field multichannel speech processing is su…

Audio-Visual Speech RecognitionAutomatic Speech RecognitionAutomatic Speech Recognition (ASR)Contrastive Learning+6

Rank-1 Constrained Multichannel Wiener Filter for Speech Recognition in Noisy Environments

2017-07-01 · Ziteng Wang, Emmanuel Vincent, Romain Serizel, Yonghong Yan

Multichannel linear filters, such as the Multichannel Wiener Filter (MWF) and the Generalized Eigenvalue (GEV) beamformer are popular signal processing techniques which can improve speech recognition performance. In this…

speech-recognitionSpeech Recognition

Joint Sound Source Separation and Speaker Recognition

2016-04-29 · Jeroen Zegers, Hugo Van hamme

Non-negative Matrix Factorization (NMF) has already been applied to learn speaker characterizations from single or non-simultaneous speech for speaker recognition applications. It is also known for its good performance i…

blind source separationSpeaker Recognition

CNN-based MultiChannel End-to-End Speech Recognition for everyday home environments

2018-11-07 · Nelson Yalta, Shinji Watanabe, Takaaki Hori, Kazuhiro Nakadai 외

Casual conversations involving multiple speakers and noises from surrounding devices are common in everyday environments, which degrades the performances of automatic speech recognition systems. These challenging charact…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Decoderspeech-recognition+1

End-to-end Multichannel Speaker-Attributed ASR: Speaker Guided Decoder and Input Feature Analysis

2023-10-16 · Can Cui, Imran Ahamad Sheikh, Mostafa Sadeghi, Emmanuel Vincent

We present an end-to-end multichannel speaker-attributed automatic speech recognition (MC-SA-ASR) system that combines a Conformer-based encoder with multi-frame crosschannel attention and a speaker-attributed Transforme…

Automatic Speech RecognitionDecoderSpeaker Identificationspeech-recognition+1