ASR data augmentation in low-resource settings using cross-lingual multi-speaker TTS and cross-lingual voice conversion
We explore cross-lingual multi-speaker speech synthesis and cross-lingual voice conversion applied to data augmentation for automatic speech recognition (ASR) systems in low/medium-resource scenarios. Through extensive experiments, we show that our approach permits the application of speech synthesis and voice conversion to improve ASR systems using only one target-language speaker during model training. We also managed to close the gap between ASR models trained with synthesized versus human speech compared to other works that use many speakers. Finally, we show that it is possible to obtain promising ASR training results with our data augmentation method using only a single real speaker in a target language.
Code (1)
Tasks
Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Data Augmentationspeech-recognitionSpeech RecognitionSpeech SynthesisVoice ConversionSimilar Papers 제목 키워드 기반
Cross-lingual Back-Parsing: Utterance Synthesis from Meaning Representation for Zero-Resource Semantic Parsing
Recent efforts have aimed to utilize multilingual pretrained language models (mPLMs) to extend semantic parsing (SP) across multiple languages without requiring extensive annotations. However, achieving zero-shot cross-l…
Cross-Lingual TransferData AugmentationSemantic ParsingZero-Shot Cross-Lingual TransferGeneralized Data Augmentation for Low-Resource Translation
Translation to or from low-resource languages LRLs poses challenges for machine translation in terms of both adequacy and fluency. Data augmentation utilizing large amounts of monolingual data is regarded as an effective…
Data AugmentationMachine TranslationTranslationUnsupervised Machine TranslationACLM: A Selective-Denoising based Generative Data Augmentation Approach for Low-Resource Complex NER
Complex Named Entity Recognition (NER) is the task of detecting linguistically complex named entities in low-context text. In this paper, we present ACLM Attention-map aware keyword selection for Conditional Language Mod…
Data AugmentationDenoisingLanguage Modellingnamed-entity-recognition+4CLASP: Few-Shot Cross-Lingual Data Augmentation for Semantic Parsing
A bottleneck to developing Semantic Parsing (SP) models is the need for a large volume of human-labeled training data. Given the complexity and cost of human annotation for SP, labeled data is often scarce, particularly …
Data AugmentationSemantic ParsingLearning Cross-lingual Mappings for Data Augmentation to Improve Low-Resource Speech Recognition
Exploiting cross-lingual resources is an effective way to compensate for data scarcity of low resource languages. Recently, a novel multilingual model fusion technique has been proposed where a model is trained to learn …
Data Augmentationspeech-recognitionSpeech RecognitionTransliteration