paper-with-me

홈 › Papers

Exploring the Impact of Data Quantity on ASR in Extremely Low-resource Languages

2024-09-13 · Yao-Fei Cheng, Li-Wei Chen, Hung-Shin Lee, Hsin-Min Wang

This study investigates the efficacy of data augmentation techniques for low-resource automatic speech recognition (ASR), focusing on two endangered Austronesian languages, Amis and Seediq. Recognizing the potential of self-supervised learning (SSL) in low-resource settings, we explore the impact of data volume on the continued pre-training of SSL models. We propose a novel data-selection scheme leveraging a multilingual corpus to augment the limited target language data. This scheme utilizes a language classifier to extract utterance embeddings and employs one-class classifiers to identify utterances phonetically and phonologically proximate to the target languages. Utterances are ranked and selected based on their decision scores, ensuring the inclusion of highly relevant data in the SSL-ASR pipeline. Our experimental results demonstrate the effectiveness of this approach, yielding substantial improvements in ASR performance for both Amis and Seediq. These findings underscore the feasibility and promise of data augmentation through cross-lingual transfer learning for low-resource language ASR.

📄 PDF Abstract BibTeX arXiv:2409.08872

Code (0)

등록된 구현이 없습니다.

Tasks

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Cross-Lingual TransferData AugmentationSelf-Supervised Learningspeech-recognitionSpeech RecognitionTransfer Learning

Similar Papers 제목 키워드 기반

Quality versus Quantity: Building Catalan-English MT Resources

2022-06-01 · SIGUL (LREC) 2022 6 · Ona de Gibert Bonet, Ksenia Kharitonova, Blanca Calvo Figueras, Jordi Armengol-Estapé 외

In this work, we make the case of quality over quantity when training a MT system for a medium-to-low-resource language pair, namely Catalan-English. We compile our training corpus out of existing resources of varying qu…

Cross-Lingual TransferTransfer LearningTranslation

Dict-NMT: Bilingual Dictionary based NMT for Extremely Low Resource Languages

2021-11-16 · ACL ARR November 2021 11 · Anonymous

Neural Machine Translation (NMT) models have been effective on large bilingual datasets. However, the existing methods and techniques show that the model's performance is highly dependent on the number of examples in tra…

Machine TranslationNMTTranslation

Dict-NMT: Bilingual Dictionary based NMT for Extremely Low Resource Languages

2022-06-09 · Nalin Kumar, Deepak Kumar, Subhankar Mishra

Neural Machine Translation (NMT) models have been effective on large bilingual datasets. However, the existing methods and techniques show that the model's performance is highly dependent on the number of examples in tra…

Machine TranslationNMTTranslation

Valuing Uncertainties in Wind Generation: An Agent-Based Optimization Approach

2022-10-06 · Daniel Shen, Marija Ilic

The increasing integration of variable renewable energy sources such as wind and solar will require new methods of managing generation uncertainty. Existing practices of uncertainty management for these resources largely…

Management

Paths for uncertainty: Exploring the intricacies of uncertainty identification for news

2018-06-01 · WS 2018 6 · Chrysoula Zerva, Sophia Ananiadou

Currently, news articles are produced, shared and consumed at an extremely rapid rate. Although their quantity is increasing, at the same time, their quality and trustworthiness is becoming fuzzier. Hence, it is importan…

Articles