paper-with-me

홈 › Papers

Enabling ASR for Low-Resource Languages: A Comprehensive Dataset Creation Approach

2024-06-03 · Ara Yeroyan, Nikolay Karpov

In recent years, automatic speech recognition (ASR) systems have significantly improved, especially in languages with a vast amount of transcribed speech data. However, ASR systems tend to perform poorly for low-resource languages with fewer resources, such as minority and regional languages. This study introduces a novel pipeline designed to generate ASR training datasets from audiobooks, which typically feature a single transcript associated with hours-long audios. The common structure of these audiobooks poses a unique challenge due to the extensive length of audio segments, whereas optimal ASR training requires segments ranging from 4 to 15 seconds. To address this, we propose a method for effectively aligning audio with its corresponding text and segmenting it into lengths suitable for ASR training. Our approach simplifies data preparation for ASR systems in low-resource languages and demonstrates its application through a case study involving the Armenian language. Our method, which is "portable" to many low-resource languages, not only mitigates the issue of data scarcity but also enhances the performance of ASR models for underrepresented languages.

📄 PDF Abstract BibTeX arXiv:2406.01446

Code (0)

등록된 구현이 없습니다.

Tasks

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognitionSpeech Recognition

Similar Papers 제목 키워드 기반

GhanaNLP Parallel Corpora: Comprehensive Multilingual Resources for Low-Resource Ghanaian Languages

2026-03-14 · Lawrence Adu Gyamfi, Paul Azunre, Stephen Edward Moore, Joel Budu 외 arxiv

Low resource languages present unique challenges for natural language processing due to the limited availability of digitized and well structured linguistic data. To address this gap, the GhanaNLP initiative has develope…

Machine Translation

Investigating an approach for low resource language dataset creation, curation and classification: Setswana and Sepedi

2020-02-18 · LREC 2020 5 · Vukosi Marivate, Tshephisho Sefara, Vongani Chabalala, Keamogetswe Makhaya 외

The recent advances in Natural Language Processing have been a boon for well-represented languages in terms of available curated data and research resources. One of the challenges for low-resourced languages is clear gui…

ClassificationData AugmentationGeneral ClassificationTopic Classification

Tibetan Language and AI: A Comprehensive Survey of Resources, Methods and Challenges

2025-10-22 · Cheng Huang, Nyima Tashi, Fan Gao, Yutong Liu 외 arxiv

Tibetan, one of the major low-resource languages in Asia, presents unique linguistic and sociocultural characteristics that pose both challenges and opportunities for AI research. Despite increasing interest in developin…

Cross-Lingual TransferMachine TranslationSpeech Recognition

Low resource language dataset creation, curation and classification: Setswana and Sepedi -- Extended Abstract

2020-03-30 · Vukosi Marivate, Tshephisho Sefara, Vongani Chabalala, Keamogetswe Makhaya 외

The recent advances in Natural Language Processing have only been a boon for well represented languages, negating research in lesser known global languages. This is in part due to the availability of curated data and res…

Data AugmentationGeneral ClassificationTopic Classification

SenWiCh: Sense-Annotation of Low-Resource Languages for WiC using Hybrid Methods

2025-05-29 · Roksana Goworek, Harpal Karlcut, Muhammad Shezad, Nijaguna Darshana 외

This paper addresses the critical need for high-quality evaluation datasets in low-resource languages to advance cross-lingual transfer. While cross-lingual transfer offers a key strategy for leveraging multilingual pret…

Cross-Lingual TransferMultilingual NLP