Pretraining by Backtranslation for End-to-end ASR in Low-Resource Settings
We explore training attention-based encoder-decoder ASR in low-resource settings. These models perform poorly when trained on small amounts of transcribed speech, in part because they depend on having sufficient target-side text to train the attention and decoder networks. In this paper we address this shortcoming by pretraining our network parameters using only text-based data and transcribed speech from other languages. We analyze the relative contributions of both sources of data. Across 3 test languages, our text-based approach resulted in a 20% average relative improvement over a text-based augmentation technique without pretraining. Using transcribed speech from nearby languages gives a further 20-30% relative reduction in character error rate.
Code (0)
등록된 구현이 없습니다.
Tasks
Data AugmentationDecoderSimilar Papers 제목 키워드 기반
Low-resource Neural Machine Translation: Benchmarking State-of-the-art Transformer for Wolof<->French
In this paper, we propose two neural machine translation (NMT) systems (French-to-Wolof and Wolof-to-French) based on sequence-to-sequence with attention and Transformer architectures. We trained our models on the parall…
BenchmarkingLow Resource Neural Machine TranslationLow-Resource Neural Machine TranslationMachine Translation+3Backtranslation in Neural Morphological Inflection
Backtranslation is a common technique for leveraging unlabeled data in low-resource scenarios in machine translation. The method is directly applicable to morphological inflection generation if unlabeled word forms are a…
Machine TranslationMorphological InflectionTranslationAdobe AMPS’s Submission for Very Low Resource Supervised Translation Task at WMT20
In this paper, we describe our systems submitted to the very low resource supervised translation task at WMT20. We participate in both translation directions for Upper Sorbian-German language pair. Our primary submission…
Machine TranslationTranslationSamsung R&D Institute Poland submission to WAT 2021 Indic Language Multilingual Task
This paper describes the submission to the WAT 2021 Indic Language Multilingual Task by Samsung R&D Institute Poland. The task covered translation between 10 Indic Languages (Bengali, Gujarati, Hindi, Kannada, Malayalam,…
Domain AdaptationKnowledge DistillationNMTTranslation+1The University of Edinburgh’s English-Tamil and English-Inuktitut Submissions to the WMT20 News Translation Task
We describe the University of Edinburgh’s submissions to the WMT20 news translation shared task for the low resource language pair English-Tamil and the mid-resource language pair English-Inuktitut. We use the neural mac…
Language ModelingLanguage ModellingMachine TranslationTranslation