Using Spoken Word Posterior Features in Neural Machine Translation
A spoken language translation (ST) system consists of at least two modules: an automatic speech recognition (ASR) system and a machine translation (MT) system. In most cases, an MT is only trained and optimized using error-free text data. If the ASR makes errors, the translation accuracy will be greatly reduced. Existing studies have shown that training MT systems with ASR parameters or word lattices can improve the translation quality. However, such an extension requires a large change in standard MT systems, resulting in a complicated model that is hard to train. In this paper, a neural sequence-to-sequence ASR is used as feature processing that is trained to produce word posterior features given spoken utterances. The resulting probabilistic features are used to train a neural MT (NMT) with only a slight modification. Experimental results reveal that the proposed method improved up to 5.8 BLEU scores with synthesized speech or 4.3 BLEU scores with the natural speech in comparison with a conventional cascaded-based ST system that translates from the 1-best ASR candidates.
Code (0)
등록된 구현이 없습니다.
Tasks
Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Machine TranslationNMTspeech-recognitionSpeech RecognitionTranslationSimilar Papers 제목 키워드 기반
Log-linear Models for Uyghur Segmentation in Spoken Language Translation
To alleviate data sparsity in spoken Uyghur machine translation, we proposed a log-linear based morphological segmentation approach. Instead of learning model only from monolingual annotated corpus, this approach optimiz…
Machine TranslationSegmentationTranslationWord AlignmentJoint ASR and MT Features for Quality Estimation in Spoken Language Translation
This paper aims to unravel the automatic quality assessment for spoken language translation (SLT). More precisely, we propose several effective estimators based on our estimation of transcription (ASR) quality, translati…
TranslationAssessing the Tolerance of Neural Machine Translation Systems Against Speech Recognition Errors
Machine translation systems are conventionally trained on textual resources that do not model phenomena that occur in spoken language. While the evaluation of neural machine translation systems on textual inputs is activ…
Automatic Speech RecognitionAutomatic Speech Recognition (ASR)DecoderMachine Translation+5CUHK System for QUESST Task of MediaEval 2014
This paper describes a spoken keyword search system developed at the Chinese University of Hong Kong (CUHK) for the query by example search on speech (QUESST) task of MediaEval 2014. This system utilizes posterior featur…
ClusteringDynamic Time WarpingKeyword SpottingData Augmentation for Sign Language Gloss Translation
Sign language translation (SLT) is often decomposed into video-to-gloss recognition and gloss-to-text translation, where a gloss is a sequence of transcribed spoken-language words in the order in which they are signed. W…
Data AugmentationLow Resource Neural Machine TranslationLow-Resource Neural Machine TranslationLow Resource NMT+4