Hybrid Attention-based Encoder-decoder Model for Efficient Language Model Adaptation
The attention-based encoder-decoder (AED) speech recognition model has been widely successful in recent years. However, the joint optimization of acoustic model and language model in end-to-end manner has created challenges for text adaptation. In particular, effective, quick and inexpensive adaptation with text input has become a primary concern for deploying AED systems in the industry. To address this issue, we propose a novel model, the hybrid attention-based encoder-decoder (HAED) speech recognition model that preserves the modularity of conventional hybrid automatic speech recognition systems. Our HAED model separates the acoustic and language models, allowing for the use of conventional text-based language model adaptation techniques. We demonstrate that the proposed HAED model yields 23% relative Word Error Rate (WER) improvements when out-of-domain text data is used for language model adaptation, with only a minor degradation in WER on a general test set compared with the conventional AED model.
Code (0)
등록된 구현이 없습니다.
Tasks
Automatic Speech RecognitionDecoderLanguage ModelingLanguage Modellingmodelspeech-recognitionSpeech RecognitionSimilar Papers 제목 키워드 기반
Enhanced Hybrid Transducer and Attention Encoder Decoder with Text Data
A joint speech and text optimization method is proposed for hybrid transducer and attention-based encoder decoder (TAED) modeling to leverage large amounts of text corpus and enhance ASR accuracy. The joint TAED (J-TAED)…
DecoderDomain AdaptationA Unified Framework for Efficient Remote Sensing Visual Question Answering: Adapting Dual, Hybrid, and Encoder-Decoder Architectures
Visual Question Answering (VQA) in the Remote Sensing (RS) domain presents unique challenges due to the high resolution, multi scale object distribution, and semantic complexity of aerial imagery. While general domain Fo…
Visual Question AnsweringMultimodal ReasoningAdvances in Joint CTC-Attention based End-to-End Speech Recognition with a Deep CNN Encoder and RNN-LM
We present a state-of-the-art end-to-end Automatic Speech Recognition (ASR) model. We learn to listen and write characters with a joint Connectionist Temporal Classification (CTC) and attention-based encoder-decoder netw…
Automatic Speech RecognitionAutomatic Speech Recognition (ASR)DecoderGeneral Classification+3Modular Hybrid Autoregressive Transducer
Text-only adaptation of a transducer model remains challenging for end-to-end speech recognition since the transducer has no clearly separated acoustic model (AM), language model (LM) or blank model. In this work, we pro…
DecoderLanguage ModelingLanguage Modellingspeech-recognition+1An Attention-LSTM Hybrid Model for the Coordinated Routing of Multiple Vehicles
Reinforcement learning has recently shown promise in learning quality solutions in a number of combinatorial optimization problems. In particular, the attention-based encoder-decoder models show high effectiveness on var…
Combinatorial OptimizationComputational EfficiencyDecoderTraveling Salesman Problem