paper-with-me

Papers

From Sinhala to Dhivehi: Cross-Lingual Transfer Learning for Low-Resource Speech Recognition

2026-07-07 · Lukmal Ilyas, Nevidu Jayatilleke arxiv

Dhivehi, the national language of the Maldives, is currently under-resourced for automatic speech recognition (ASR) and other NLP tasks. This study investigates whether cross-lingual transfer learning from Sinhala, a linguistically related, relatively well-resourced Insular Indo-Aryan language, can improve Dhivehi ASR. We conduct seventeen experiments across five transfer learning paradigms: Dhivehi-only baselines, sequential fine-tuning, multilingual fine-tuning, continual pre-training, and a control using Turkish as an unrelated language. The strongest system, continual pre-training on Sinhala followed by fine-tuning on Dhivehi with KenLM, achieves 12.89% WER and 2.70% CER, outperforming the Dhivehi-only baseline by 13.50% WER and 3.02% CER. However, the adaptation strategy and decoding configuration are equally critical for a successful transfer learning experiment. We conduct seventeen controlled experiments spanning five transfer learning paradigms: Dhivehi-only baselines, sequential fine-tuning, multilingual fine-tuning, continual pre-training, and a control experiment using Turkish as an unrelated language. The strongest system, continual pre-training on Sinhala followed by fine-tuning on Dhivehi with KenLM, achieves 12.89% WER and 2.70% CER, outperforming the Dhivehi-only baseline by 13.50% WER and 3.02% CER. The Turkish control experiment confirms that observed improvements stem from linguistic relatedness; adaptation strategy and decoding configuration are also critical.

📄 PDF Abstract BibTeX arXiv:2607.06289

Code (0)

등록된 구현이 없습니다.

Tasks

Cross-Lingual TransferSpeech RecognitionTransfer Learning

Similar Papers 제목 키워드 기반

Detecting Urgency Status of Crisis Tweets: A Transfer Learning Approach for Low Resource Languages

2020-12-01 · COLING 2020 8 · Efsun Sarioglu Kayi, Linyong Nan, Bohan Qu, Mona Diab 외

We release an urgency dataset that consists of English tweets relating to natural crises, along with annotations of their corresponding urgency status. Additionally, we release evaluation datasets for two low-resource la…

Transfer LearningXLM-R

Sinhala-English Word Embedding Alignment: Introducing Datasets and Benchmark for a Low Resource Language

2023-11-17 · Kasun Wickramasinghe, Nisansa de Silva

Since their inception, embeddings have become a primary ingredient in many flavours of Natural Language Processing (NLP) tasks supplanting earlier types of representation. Even though multilingual embeddings have been us…

Exploiting Parallel Corpora to Improve Multilingual Embedding based Document and Sentence Alignment

2021-06-12 · Dilan Sachintha, Lakmali Piyarathna, Charith Rajitha, Surangika Ranathunga

Multilingual sentence representations pose a great advantage for low-resource languages that do not have enough data to build monolingual models on their own. These multilingual sentence representations have been separat…

Sentence

Utilizing Multilingual Encoders to Improve Large Language Models for Low-Resource Languages

2025-08-12 · Imalsha Puranegedara, Themira Chathumina, Nisal Ranathunga, Nisansa de Silva 외 arxiv

Large Language Models (LLMs) excel in English, but their performance degrades significantly on low-resource languages (LRLs) due to English-centric training. While methods like LangBridge align LLMs with multilingual enc…

News Classification

Neural Machine Translation for Sinhala-English Code-Mixed Text

2021-09-01 · RANLP 2021 9 · Archchana Kugathasan, Sagara Sumathipala

Code-mixing has become a moving method of communication among multilingual speakers. Most of the social media content of the multilingual societies are written in code-mixed text. However, most of the current translation…

DecoderMachine TranslationNMTTranslation