paper-with-me

Papers

Multi-encoder multi-resolution framework for end-to-end speech recognition

2018-11-12 · Ruizhi Li, Xiaofei Wang, Sri Harish Mallidi, Takaaki Hori, Shinji Watanabe, Hynek Hermansky

Attention-based methods and Connectionist Temporal Classification (CTC) network have been promising research directions for end-to-end Automatic Speech Recognition (ASR). The joint CTC/Attention model has achieved great success by utilizing both architectures during multi-task training and joint decoding. In this work, we present a novel Multi-Encoder Multi-Resolution (MEMR) framework based on the joint CTC/Attention model. Two heterogeneous encoders with different architectures, temporal resolutions and separate CTC networks work in parallel to extract complimentary acoustic information. A hierarchical attention mechanism is then used to combine the encoder-level information. To demonstrate the effectiveness of the proposed model, experiments are conducted on Wall Street Journal (WSJ) and CHiME-4, resulting in relative Word Error Rate (WER) reduction of 18.0-32.1%. Moreover, the proposed MEMR model achieves 3.6% WER in the WSJ eval92 test set, which is the best WER reported for an end-to-end system on this benchmark.

📄 PDF Abstract BibTeX arXiv:1811.04897

Code (0)

등록된 구현이 없습니다.

Tasks

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognitionSpeech Recognition

Similar Papers 제목 키워드 기반

SAMU-XLSR: Semantically-Aligned Multimodal Utterance-level Cross-Lingual Speech Representation

2022-05-17 · Sameer Khurana, Antoine Laurent, James Glass

We propose the SAMU-XLSR: Semantically-Aligned Multimodal Utterance-level Cross-Lingual Speech Representation learning framework. Unlike previous works on speech representation learning, which learns multilingual context…

Representation LearningRetrievalSentenceSentence Embedding+6

A Multiscale Autoencoder (MSAE) Framework for End-to-End Neural Network Speech Enhancement

2023-09-21 · Bengt J. Borgstrom, Michael S. Brandstein

Neural network approaches to single-channel speech enhancement have received much recent attention. In particular, mask-based architectures have achieved significant performance improvements over conventional methods. Th…

Automatic Speech RecognitionSpeech Enhancementspeech-recognitionSpeech Recognition

Multi-Stream End-to-End Speech Recognition

2019-06-17 · Ruizhi Li, Xiaofei Wang, Sri Harish Mallidi, Shinji Watanabe 외

Attention-based methods and Connectionist Temporal Classification (CTC) network have been promising research directions for end-to-end (E2E) Automatic Speech Recognition (ASR). The joint CTC/Attention model has achieved …

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognitionSpeech Recognition

A Multi-level Acoustic Feature Extraction Framework for Transformer Based End-to-End Speech Recognition

2021-08-18 · Jin Li, Rongfeng Su, Xurong Xie, Nan Yan 외

Transformer based end-to-end modelling approaches with multiple stream inputs have been achieved great success in various automatic speech recognition (ASR) tasks. An important issue associated with such approaches is th…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)DecoderDiversity+3

Hierarchical Multi-Grained Generative Model for Expressive Speech Synthesis

2020-09-17 · Yukiya Hono, Kazuna Tsuboi, Kei Sawada, Kei Hashimoto 외

This paper proposes a hierarchical generative model with a multi-grained latent variable to synthesize expressive speech. In recent years, fine-grained latent variables are introduced into the text-to-speech synthesis th…

Expressive Speech SynthesisSpeech Synthesistext-to-speechText to Speech+1